Practical guide
Does ChatGPT Train on My Data? How to Read AI Vendor Terms
How to find out what an AI vendor actually does with your business data — which sections of the terms matter, the phrases that signal training and retention, and what to do when the answer is unclear.
Key takeaways
- The answer depends on which tier you are on, not which product. Several major consumer tiers default to using your conversations for training and at least one does not — which is exactly why you check your own. Business and API tiers generally exclude it contractually.
- Four sections carry the real answers: how inputs are used for model improvement, retention periods, subprocessors, and confidentiality. Everything else is usually boilerplate.
- Watch for "to improve our services" — it is broad enough to include model training, and it is the most common way training permission is granted without the word "training" appearing.
- Check the account settings too. A contract that permits opting out means nothing until someone opts out on every account.
- If you cannot determine the answer in fifteen minutes of reading, treat it as unclear and keep sensitive data out until the vendor confirms in writing.
"Does ChatGPT train on my data?" is the most common AI question small business owners ask, and the honest answer is that it depends on which account you are using — which is exactly why the question keeps getting asked.
The same product can behave differently on a free login, a team plan, and an API key. So rather than answering for one vendor at one moment, this walks through how to get the answer yourself, for any AI tool, in about fifteen minutes.
Why the answer changes by tier
AI vendors generally treat consumer and business relationships differently.
On consumer tiers, the product is often improved using what users type. Several of the largest chat interfaces default to this and give you a settings toggle to turn it off. But it is not universal — at least one major provider treats model training on consumer conversations as something you have to switch on, not off.
That inconsistency is the actual lesson. Two people using two well-known chat tools on comparable paid consumer plans can be under opposite defaults, so a rule of thumb about "consumer tiers" will eventually be wrong about the specific tool your team is using. Check the tool in front of you.
On business, team, enterprise, and API tiers, vendors typically commit contractually that customer inputs are not used to train their models. This is a competitive necessity — companies would not adopt the tools otherwise.
The practical consequence: an employee using a personal free account for work may be operating under completely different data terms than the same person using the company account. That gap is why the approved-tool list in ChatGPT rules for employees specifies accounts, not just tools.
The four sections that actually matter
Vendor terms are long. Almost all of the useful information is in four places.
1. How your inputs are used
Search the terms and privacy policy for: train, training, model improvement, improve our services, machine learning.
What you want to find is an explicit statement about customer content and model training. Good language is unambiguous — something to the effect that the vendor does not use business customer content to train its models.
The phrase to be careful with is "to improve our services." It sounds administrative and is broad enough to cover model training without ever using the word. If that is the only relevant language you can find, you have not found an exclusion.
2. Retention
Search for: retain, retention, delete, deletion, storage period.
You are looking for a stated duration and a deletion mechanism. Three questions:
- How long are inputs kept by default?
- Can you delete them, and does deleting a conversation delete the underlying record?
- Is there a separate, longer retention for abuse monitoring or legal holds? A 30-day trust-and-safety retention alongside a "we do not train on your data" commitment is normal and worth knowing about.
An absent retention period is the most common gap. Silence is not "we delete it promptly."
3. Subprocessors
Search for: subprocessor, sub-processor, third party, service providers.
Many AI products are wrappers around another company's model. The list tells you who else holds your data. What matters for a small business is whether the chain is disclosed at all — if it is not, you cannot answer a customer asking where their information went.
4. Confidentiality and ownership
Search for: confidential, your content, ownership, license.
You are checking two things: that you retain ownership of what you put in, and that the license you grant the vendor is limited to operating the service rather than being broad and perpetual.
Terms are only half the answer — check the settings
A contract that allows opting out of training does nothing until somebody opts out.
For each approved tool, open the account settings and confirm the data controls are configured the way you assumed. Then note who checked and when. This takes two minutes per tool and is the single highest-value step in this entire process, because it is where the paper answer and the real answer most often diverge.
Do this per account, not per tool. Settings are usually account-scoped, so a correctly configured company workspace tells you nothing about the personal login someone uses at home.
When the answer is unclear
Sometimes fifteen minutes of reading produces no clear answer. That is a legitimate finding, and the response is straightforward:
- Record it as unclear rather than leaving it blank. An unanswered question that is written down gets asked eventually; one that is not written down does not.
- Email the vendor and ask directly whether inputs are used for model training and how long they are retained. Business-focused vendors answer this routinely.
- Until they answer, keep customer PII, financial records, and anything confidential out of the tool. Internal drafts and public information are usually fine in the meantime.
That last step is the important one. "We are still checking" is a workable position; "we did not check and used it for everything" is not.
Fitting this into a decision
Reading the terms is one input to approving a tool. The full set of questions — including what happens when the output is wrong, which the terms will never tell you — is in the AI vendor review checklist. The workflow around who reviews and records it is in the AI tool approval process.
Whatever you conclude, the conclusion belongs in a written rule your team can actually find. How to write an AI usage policy covers the document that holds it.
Before you paste, check the prompt
Terms tell you what a vendor may do with your data. They do not stop sensitive data going in. The Prompt Safety Checker scans a draft prompt for credentials, personal information, and internal financial data before you send it — running entirely in your browser, with nothing transmitted anywhere.
For a vendor review form, a data input rules checklist, and employee guidance you can hand out as-is, see the Starter Pack.
This article provides practical operational guidance for AI usage management. It is not legal advice and does not guarantee regulatory compliance. Vendor terms change and vary by tier and region — verify the current terms for your own account, and review them with qualified professionals where appropriate.