Blog

The hidden data layer in AI terms: what Usage Data really means

Major AI providers quietly claim ownership and training rights over Usage Data outside visible customer data. What this means for high-trust workflows.

· Victor Angelier

In July 2026 the Grove Foundation published an analysis that exposes a specific and often overlooked detail in the terms of use of major AI providers. According to The Grove Foundation Finds AI Platforms Quietly Claim Ownership Over Usage Data, some providers left the textual definition of 'Usage Data' unchanged over a fourteen-month period, but quietly expanded the surrounding clauses: an explicit ownership claim ('all right, title, and interest') was added, and training on Usage Data shifted from opt-in to opt-out. Crucially, in those terms Usage Data is often contractually placed outside the 'Customer Data' category — meaning that deletion rights, training opt-outs and zero-retention promises do not apply to it.

This is precisely the type of risk that does not appear in marketing material, but in the fine-grained legal definitions that determine which data counts as the provider's property. For professionals working with confidential or high-trust information, this is relevant: the protection you get on one data layer does not automatically apply to the other.

Two data layers that do not receive the same protection

The separation between 'customer data' and 'usage data' becomes tangible in ConductAtlas's analyses of OpenAI. In OpenAI Enterprise Privacy it states that, for enterprise and API customers, OpenAI promises not to carry out model training on business data unless there is opt-in, with a default retention of 30 days for API inputs and outputs and even zero-retention options. Those assurances, however, are textually focused on content classified as customer data. The precise delineation of Usage Data remains outside that regime, leaving room for a separate category that is not explicitly covered by the same protection rules.

The contrast becomes sharper in OpenAI Privacy Policy. There, ConductAtlas makes it explicit that user-submitted content — prompts, files, media — may by default be used to train models in consumer products, unless the user actively opts out. Moreover, trained, de-identified content cannot be reversed after a deletion request. Broad data categories, from chat content to advertising and partner data, are captured under a single umbrella term. In other words: alongside its business no-training promises, the same provider maintains a consumer policy in which usage and content data are indeed trainable by default, and in which opt-outs do not apply to data that has already been absorbed.

Many compliance teams look solely at the first layer — the visible customer data with short retention and no-training promises — while the contract text deliberately places a second, more broadly defined layer outside 'Customer Data'.

Grey areas around risky applications

Beyond the data layers, terms also create ambiguity about responsibility. The academic study Regulatory gray areas of LLM Terms analyses the terms of several major providers and identifies 'regulatory gray areas'. On the one hand, terms explicitly prohibit sensitive applications such as criminal justice, large-scale profiling and emotion inference, but on the other hand they leave broad clauses on data collection and risky use to self-classification by the user. Through clauses on prohibited professional advice and high-risk healthcare use, providers shift part of the liability risk back to customers, while the definition of exactly what falls under 'high risk' remains ambiguous.

For lawyers, doctors and financial institutions, this means they bear the responsibility for use in those grey areas, while their contractual rights to inspect training and usage data are limited.

Why Usage Data becomes technically indispensable

The tension between visible promises and hidden processing recurs in new safety features. According to Bloomberg, since 19 August 2026 OpenAI has been testing a new safety processing for paying tool customers, intended to recognise risk patterns across multiple interactions. It is emphasised that certain customers receive zero data retention and that prompts and responses are not stored or accessed, while that protection does not necessarily apply to all paid customers or all processing. At the same time, additional safety processing for other segments relies precisely on centralised pattern analysis across interactions — which implies that Usage Data in that context is retained and searched, even where the marketing language emphasises zero retention for specific endpoints.

This layer is therefore technically indispensable for risk management, but often poorly visible contractually.

From fine print to visible data classes

The common thread through these sources: the real contract risks lie not in overt privacy promises, but in how Usage Data is defined, claimed and used. Anyone working with sensitive information is therefore better off structuring their AI risk analysis along data classes. Record, per provider, which categories fall under customer control (no-training, short retention, deletion rights) and which are tacitly classified as the provider's property, how training rights and retention periods differ, and where rights to deletion, opt-out and audit are absent.

In that context, a verification layer such as IamVera.ai is relevant. Vera is not a chatbot and not its own language model, but a privacy-focused verification layer for professionals in high-trust environments. Vera can route a task through selected independent AI models and expose verification steps, corrections, disagreements and sources for inspection. This supports review and oversight, but does not replace the professional's own judgement of correctness.

For the data layer itself, the Semantic Privacy Shield is designed to replace sensitive document values with synthetic, session-only equivalents on EU infrastructure before AI processing. The AI chain analyses the synthetic version; the original values can then be restored locally. The architecture is designed so that only anonymised content is sent onward, and the workflow is fail-closed: if the privacy check fails, the document is not sent. Uploaded PDFs are processed temporarily for the active run and are not stored permanently; metadata may be retained for session history.

Such measures are no guarantee of flawless anonymisation or full GDPR compliance, and the final professional judgement always remains with the user. But they can help to explicitly factor in the hidden Usage layer in the assessment — before a contract is signed or an AI workflow goes live.

← All articles