Why You Will Probably Never Know Which of Your Data Went Into AI Training
Set Trending Topics as a preferred source on Google. Since last year, the EU has had a rule on the books that sounds, at first glance, like the end of the secrecy: anyone placing a general purpose AI model on t...
- 6 min read
Set Trending Topics as a preferred source on Google.
Since last year, the EU has had a rule on the books that sounds, at first glance, like the end of the secrecy: anyone placing a general purpose AI model on the European market has to publish a summary of the content it was trained on. The obligation sits in Article 53 of the AI Act, and the format is prescribed by the European Commission through a mandatory template.
Anyone hoping to finally look up whether their own blog, their own book or their own photos ended up inside ChatGPT, Claude or Gemini will be disappointed, though. The Commission’s FAQ on the template spells out fairly plainly where transparency stops.
What providers have to disclose
The list is not short. Providers have to state which model and which model versions are covered, what types of content went into training (text, image, audio, video), the broad size range of the data and its general characteristics. Publicly available datasets have to be listed. For licensed or privately sourced datasets, a description without the commercial details is enough. For user data, providers have to disclose which services and which modalities it came from, while personal information stays out. Synthetic data has to be described as well, along with the measures taken against illegal content and anything that affects the exercise of copyright.
The most interesting item is web scraping. Here providers have to name the crawlers they used, state the periods over which collection took place and describe what was crawled. And they have to disclose the most important domains the data came from.
And this is exactly where transparency ends
The number attached to that requirement is the heart of the matter: what has to be disclosed is the top ten percent of all domains, measured by volume of data. For small and medium sized enterprises, it is the top five percent or 1,000 domains. Everything below that, which for models learning from billions of web pages is the overwhelming majority, stays invisible. If your website is not among the internet’s major data suppliers, it will not show up in any of these summaries, even if it was scraped in full.
On top of that come further carve outs the FAQ names explicitly. Individual works do not have to be listed. Neither do individual URLs. Datasets below the thresholds drop out. Trade secrets and confidential business information have to be weighed against the interest in transparency, which in practice means providers decide for themselves where to draw that line. And anyone who, despite diligent efforts, can no longer reconstruct certain information merely has to explain the gap, not close it.
The summary is therefore a catalogue of categories, not a search engine. It tells you that a model processed, say, Common Crawl, Wikipedia and a number of large news sites. It does not tell you whether your text was among them.
The timeline pushes the answer further out
The deadlines do their part to keep the question open. Models that came to market before the obligation took effect, and that includes many of the systems in everyday use today, have until August of next year to file their summary. After that, updates are due every six months, or sooner if substantial new training data is added.
The GPAI obligations have only been enforceable since this summer. Since then, fines of up to three percent of global annual turnover or 15 million euros, whichever is higher, are on the table. Article 53 itself was left untouched by the Digital Omnibus, with which the Commission softened other parts of the AI Act.
Japan goes one step further, but without teeth
The comparison with Japan is instructive. A government expert panel there recently approved a code for providers of generative AI, as reported by the Japan Times. Its second principle would have companies answer requests from rights holders about whether specific web pages are contained in the training data.
That is precisely what the EU template leaves out. The catch: the Japanese code is not legally binding. Companies can opt for a comply or explain approach and publicly justify why they will not take part.
What you can do instead
Anyone who wants to know whether their own content sits inside a model still has to rely on indirect methods. Researchers work with memorisation tests, prompting models into reproducing longer passages verbatim. One prominent example was the finding that Llama had effectively memorised large parts of Harry Potter. Tests like these mainly work for texts that appear very frequently, not for an average blog post.
The second route runs through the courts. In the US, numerous copyright cases around AI training are underway in which the origin of the data has to be disclosed during discovery. What surfaces there has so far been considerably more detailed than anything providers publish voluntarily or on the basis of the template.
That leaves the view forward: anyone who does not want future models trained on their content can invoke the rights reservation for text and data mining, technically via robots.txt or machine readable usage restrictions. That only works for future crawling runs, though. Whatever is already baked into the weights of a trained model cannot be pulled back out.
Opt-outs at Instagram, Facebook or LinkedIn
At the large platforms themselves there is at least a documented way out, because they train their AI models on the content of their own users. Meta processes public posts, comments and profile information from adult users in the EU to train its AI, relying on “legitimate interest” under the GDPR. That is exactly why there is a right to object. The route runs through dedicated forms, one for Facebook and one for Instagram, also reachable through the privacy policy in the settings by searching for “objection” there. According to Meta, the objection applies permanently and does not have to be repeated; for linked accounts, one submission covers them all. For WhatsApp there is a precautionary form, even though Meta says it does not currently use WhatsApp data for training.
LinkedIn has also been using the data of European members to train its generative AI since last year, with the accounts of minors and private messages excluded. The opt-out sits in the settings under “Data privacy” in the section on how LinkedIn uses your data; the relevant toggle is called “Data for Generative AI Improvement”. It is switched on by default and can be turned off with a single click, as we described here.
The same limitation applies in both cases as with the rights reservation: an objection works going forward. No form retrieves what has already flowed into a model. And how much of that is your material is something the official summary, as things stand today, will not tell you.