Data Governance & Lineage advanced 8 min read 7 flashcards

Consent and Purpose Limitation for Training Data

Why data collected for one purpose cannot simply be reused to train a model under GDPR, how the Article 6(4) compatibility test and the legitimate-interest route work in practice, what EDPB Opinion 28/2024 and the courts have said, and where the law stands as of September 2026.

A company holds five years of support chats, collected to resolve tickets, and a team wants to fine-tune a support model on them. The data is in the warehouse, access controls are fine, and the distribution is exactly right. Under GDPR none of that settles whether the training is lawful, because the question is not who may read the data but what it was collected for.

Article 5(1)(b) requires personal data to be "collected for specified, explicit and legitimate purposes and not further processed in a manner that is incompatible with those purposes", with archiving, scientific or historical research and statistics carved out under Article 89 safeguards (Regulation (EU) 2016/679, Art. 5). Training is further processing, and every training run needs two things: a purpose compatible with collection, and a lawful basis under Article 6. This concept complements Provenance for Training Corpora, which records where data came from, and Retention, Deletion and the Right to Erasure, which handles what happens when a person leaves.

The compatibility test

When further processing is not based on consent or a specific law, Article 6(4) directs the controller to weigh, among other things: the link between the original and new purposes; the context of collection, especially the relationship with the data subject; the nature of the data, especially Article 9 special categories and criminal data; the possible consequences for data subjects; and safeguards such as encryption or pseudonymisation (GDPR Art. 6(4)).

Applied to the support chats: the link is moderate (improving support is adjacent to resolving tickets), customers expected a human conversation, chats often contain health and financial disclosures, and a model that memorises a chat could repeat it to another customer. Pseudonymisation and filtering move the balance; they do not make it automatic.

Consent (Art. 6(1)(a)) looks cleanest and behaves worst at training scale. It must be freely given, specific and informed, Article 7(3) lets it be withdrawn at any time, and withdrawal after training leaves the person's influence in the weights, which is the unsolved problem of machine unlearning.

Legitimate interest (Art. 6(1)(f)) permits processing "necessary for the purposes of the legitimate interests pursued by the controller", unless overridden by the data subject's interests or rights. The EDPB's Opinion 28/2024, adopted in December 2024 at the Irish supervisory authority's request, confirmed that it can be a basis for developing and deploying AI models, subject to a three-step assessment: identify a legitimate interest, show the processing is necessary for it, and balance it against data subjects' rights, including their reasonable expectations (EDPB, 2024, Opinion 28/2024). It also held that a model trained on personal data is anonymous only if extracting personal data, directly or through queries, is insignificant, and that unlawful processing during development can affect the lawfulness of the model's later use. The CJEU had already accepted that a purely commercial interest can be legitimate, while insisting on strict necessity and balancing (CJEU, C-621/22, Koninklijke Nederlandse Lawn Tennisbond, 4 October 2024).

Legitimate interest brings an Article 21 right to object, so the engineering consequence is the same as for consent: training jobs must read an objection register at snapshot time, and the snapshot must be recorded.

What regulators and courts have done

The picture is unsettled, and the disagreement is real. In May 2025 the Irish Data Protection Commission reviewed Meta's revised plan to train on EU users' public content under legitimate interest, with an objection mechanism, and issued recommendations rather than a prohibition; Meta's planned start date was 27 May. The Higher Regional Court of Cologne refused a consumer group's urgent injunction against the plan (OLG Köln, 15 UKl 2/25, 23 May 2025), a summary proceeding, not a final ruling. Italy's Garante took the opposite line, fining OpenAI €15 million in part for training without an adequate legal basis; on 18 March 2026 the Court of Rome annulled that decision because the Garante lacked competence once OpenAI had an Irish main establishment, and it did not rule on whether the training was lawful. Outside the EU, the US Federal Trade Commission's 2021 order against Everalbum required deletion not only of improperly obtained face data but of models and algorithms developed from it (FTC, 2021, FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology).

As of September 2026, the Commission's Digital Omnibus proposal of 19 November 2025 would add an Article 88c confirming legitimate interest for AI development and operation, with enhanced safeguards and an unconditional right to object, and a new Article 9 exception for residual special-category data in training sets. It is not law: the Council's mandate vote planned for June 2026 was cancelled for lack of agreement, and the Parliament committee report had drawn more than 1,750 amendments (European Parliament Legislative Train, The Digital Omnibus Regulation Proposal). The separate AI Omnibus amending the AI Act was adopted; the GDPR changes were not.

When it breaks

The fine arithmetic is not the main risk. Article 83(5) caps fines for breaching the basic principles at €20 million or 4% of worldwide annual turnover, whichever is higher; for a firm with €2 billion turnover that is an €80 million ceiling. A deletion order covering the model, as in Everalbum, can cost more than the fine.

Broad purpose statements fail the specificity test. "To improve our services" names no purpose a data subject could anticipate.

Objections arrive after the snapshot. A person who objects after training has a right that retraining schedules, not policies, determine how quickly you honour. Record which objection-register version each run used, or you cannot show compliance.

Scraped data breaks the notification model. Transparency duties and a right to object presume you can reach people. For web-scale corpora you mostly cannot, which is why commentators call an opt-out-based Article 88c unworkable for scraped data.

Check yourself

7 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track