Models don't retain your words, but they can reproduce them under the right conditions. Part 3 explores what research shows about data encoded in model weights, and why GDPR cannot yet protect you from what you've already shared.
Welcome to the final instalment of Who Owns Your Thoughts? If you somehow missed part 1 and part 2 I recommend you go back and read them first to get full context - kind of like a chatbot.
Here is a summary of what we now know:
AI models do not retain anything from your conversations. The surrounding AI systems in the application layer do store your conversation data - actively indexing and distilling it to construct the experience of memory. AI vendors actively collect and use your conversations for training with your (more or less forced) consent. The question which remains is, if there is any part of your initial trail of thoughts that can be traced in these models - and if so, who then owns that data?
Harry Potter and the Frozen Weights
LLMs are too large and too complex for any human to inspect in detail what is actually encoded in the model weights. It is well established that new models regularly surprise even their own creators with emergent capabilities - the ability to speak languages never explicitly included in training data, for instance. Since we cannot know for certain what a model contains, a significant part of AI safety research involves probing and testing for unexpected behaviour.
This is how a team of researchers from Google DeepMind, in 2023, found a way to recover 10,000 unique memorised training examples - including real email addresses and phone numbers, verbatim paragraphs from books and poems, URLs, and unique identifiers - simply by prompting the model in a particular way. The finding demonstrated that what a model has seen during training is not fully abstracted. Under the right conditions, it can be recovered.
2023 already feels distant in AI terms. But as recently as March 2026, a team of Stanford researchers published a paper documenting how they were able to prompt Gemini, Grok, and Claude Sonnet into reproducing text from Harry Potter and the Philosopher's Stone - simply by typing the book's opening sentence and asking the model to continue. The result was a 98.5% accurate reproduction of the full text. This confirms that LLMs are extraordinary prediction machines, but also that they store data in ways that make it retrievable under the right conditions.
So it turns out AI models do encode data during training in ways that can, under some conditions, be retrieved. Scaling laws suggest that larger future models will have an even greater capacity for this kind of implicit retention. These findings raise serious questions about whether current security constraints - at the model and system levels - are sufficient to prevent training data from being extracted, and they contradict the claim, made by several providers, that models do not memorise but merely learn statistical representations.
Your Thoughts Are Still Yours. Aren't They?
If it has not already happened, the thoughts recorded in your conversations over the years can indeed be used by the labs to shape future models. In that sense, your thoughts are still yours - just no longer exclusively.
The economic logic of this exchange is straightforward and not inherently sinister, but it is rarely stated plainly. Users receive a powerful, free or low-cost service. In exchange, their conversations become training signal that improves the model for future users and strengthens the company's commercial position. This is the same exchange model as social media, email, and search - but with AI this data is more personal, user understanding is lower, and the feedback loop (better models attract more users, who generate more training data) accelerates faster than public awareness can keep pace with.
What makes AI chat different from every previous data collection technology is not the quantity of data collected - it is the quality. Social media captured what you chose to share publicly. Email captured functional communications. Search queries captured intent signals. AI chat captures something closer to actual thought: the unresolved questions, the half-formed strategies, the personal doubts you share with a system that feels, behaviourally if not architecturally, as though it understands you.
And that data clearly has value - not just to the companies, but to you. It contains traces of who you are: what you care about, what you cook for dinner, your health concerns, your professional dilemmas, your private uncertainties. Even if the labs have a legal right to train on it, at a minimum it seems reasonable that you should be able to save it and take it with you.
But that is precisely the problem. Your conversations are siloed within walled gardens. Anthropic holds your Claude history. OpenAI holds your ChatGPT history. Google holds your Gemini history. There is no portability standard, no common format, and no practical mechanism for transferring your conversational history from one provider to another. You can export a raw data dump, and data regulations like GDPR grant you formal portability rights - but a folder of raw JSON files is not a functional equivalent of continuing an ongoing relationship with a new provider. The cognitive infrastructure that makes these tools useful - the accumulated context, the personalisation, the memory - cannot travel with you. The companies are building moats out of your memories.
The Law Is Catching Up. It Has Not Arrived.
European users have formal rights over their personal data. Under GDPR Article 17, you have the right to erasure. Under Article 15, you have the right to access what is held about you. Under Article 20, you have the right to portability.
The problem is that all of these rights were designed with databases in mind - structured records that can be found, accessed, and deleted. They collide with a fundamental technical obstacle when applied to trained models:
GDPR does not distinguish between a classic database and a trained model. This creates genuine uncertainty, because even learned representations can be personal if they are traceable. An AI model learns by processing training data. What results are weight matrices and vector representations that bear no resemblance to structured datasets - and cannot simply be deleted.
The EU AI Act introduces additional obligations targeting AI systems specifically. Under Article 53, providers of general-purpose AI models must publish a structured summary of training data, including whether user-generated interaction data was used. But once that data is encoded in a model, there is no mechanism for removal. The model cannot unlearn, and retraining from scratch is economically infeasible - a reality the AI Act itself acknowledges.
So while EU legislation points in the right direction, there is currently no enforceable remedy once your words have contributed to shaping a model's behaviour. Even in the diffuse, statistical sense the labs describe, there is no surgical removal available.
Bottom Line
AI models cannot remember what you tell them - but AI systems keep records of your thoughts, partly for your benefit and partly for theirs.
In many ways, your thoughts remain yours. You retain legal ownership of what you express. But the data governance landscape is heavily skewed towards corporate interests. By accepting the terms of service, you granted broad licences to the frontier labs. They have stored your conversations, potentially had humans review them, used them to shape model behaviour, and encoded something of them into weight matrices that cannot be fully unwound. Platform owners have a strong commercial interest in keeping you on their platform, and there is currently no standard for transferring your data and accumulated memory to a competitor in any meaningful, functional sense.
It is possible to extract your data - but only if you are technically minded. The law on all of these matters is catching up, but has not yet arrived. And the terms that govern all of it were presented to you in a pop-up you scrolled past.
That concludes our deprive into AI memory and data privacy. I hope you learned something new along the way and now understand a little better where your thoughts travel, when you send them off into AI systems.
Lars Harder
Writing on sovereign AI, digital identity, and what it means to remain human in an era of algorithmic culture.
// more reading
