The Educational Data Ghost: AI, Learning Platforms, and European Sovereignty

George Ellis
11 Min Read

Written by Jasmin B. Cowin, Ed. D.

European Digital Sovereignty

European digital sovereignty is usually framed as a question of control: who governs the data, infrastructure, standards, and jurisdiction under which digital life operates. The European Commission now treats that question as industrial policy. Its AI Continent Action Plan of April 2025 expands the network of EuroHPC AI Factories, commits the twenty billion euro InvestAI facility to as many as five AI gigafactories, each planned around more than 100,000 advanced AI processors, and proposes a Cloud and AI Development Act intended to at least triple EU data centre capacity within five to seven years (European Commission, 2025a). In this formulation, sovereignty becomes less an abstract legal principle than a material project of building, locating, and securing the computational capacity on which AI development depends.

Infrastructure and Metadata

This framing operates primarily at the level of states, institutions, and infrastructure. I propose a companion concept that shifts the analysis to the level of the person: the data ghost. By this, I mean the residual presence of a person, institution, or community within digital systems after the original interaction has ended. Such presence may persist as metadata, behavioral patterns, biometric templates, voiceprints, inferred emotional or cognitive profiles, embeddings, training data, risk scores, or algorithmic predictions. The concept of the data ghost exposes a central limit of sovereignty claims: a member state, university, or individual may formally control data, while traces of that data continue to circulate, be recombined, and support future inference. Digital sovereignty is therefore not only a question of control over digital systems. It is also a question of whether such control remains meaningful after data has been captured, copied, inferred, or embedded into AI models. On this view, sovereignty cannot refer only to the ownership of servers or the national control of platforms. It should also include control over the afterlife of data, including retention, deletion, provenance, consent, auditability, model training, and the right not to be indefinitely reconstructed from one’s own traces.

Regulations vs Reality

The EU AI Act, Regulation (EU) 2024/1689, is the first legal instrument to come close to addressing the data ghost directly, even though it never names it as such. Under Article 53(1)(d), every provider of a general-purpose AI model placed on the Union market must publish a sufficiently detailed summary of the content used for training, using the mandatory template issued by the European AI Office on 24 July 2025. Because this obligation also extends to open-source models, requires summaries to be updated at least every six months, and may, from 2 August 2026, be enforced by the AI Office through fines of up to fifteen million euros or three percent of global annual turnover, it gives legal form to a problem that had previously remained largely infrastructural and epistemic (European Commission, 2025b). In effect, the provision turns provenance into a regulatory question: it asks providers to account, at least in aggregate, for the traces that have entered the model. Yet this recognition remains partial, since the template avoids work-by-work disclosure and depends instead on narrative and aggregated reporting, thereby acknowledging the ghost as a population rather than as a person. So “the ghost as a population” means that the regulation acknowledges that many people’s traces are within the model, but it does not make those traces visible at the level of the individual.

Individuals and Personal Data Residue

The individual returns more clearly in data protection law, particularly in Opinion 28/2024, adopted on 17 December 2024, where the European Data Protection Board rejected the assumption that AI models trained on personal data can be treated as automatically anonymous. For the EDPB, anonymity must be assessed case by case and can be established only where the likelihood of extracting personal data from the model, whether directly or through queries, is insignificant for every affected individual, taking into account all means reasonably likely to be used (EDPB, 2024). The significance of this position lies in its treatment of trained weights as possible sites of identifiable residue, including residue that may be recovered through techniques such as membership inference or model inversion. It also brings Article 17 of the GDPR, the right to erasure, into contact with a difficult technical reality: while deleting a record from a database is relatively straightforward, removing a person’s contribution from a trained model is not, and machine unlearning remains an unsettled research problem. The decisive test for European digital sovereignty may therefore lie less in the regulation of collection than in the governance of what data becomes after it has been absorbed, transformed, and made difficult to retrieve.

EU Inference Providers and Data Ghosts

This is why the emerging class of EU inference providers matters, and why the term requires a precise definition. Rather than training frontier models, an inference provider hosts existing models, often open-weight models, on European infrastructure and serves them through an API, making the governance of prompts, outputs, and logs a question of where computation occurs, which legal regime applies, and how far foreign access claims, including those associated with the United States extraterritorial CLOUD Act, can reach. Inference is where ghosts are actively generated: every query produces new inferences about the person behind it. Providers in this space include Scaleway’s Generative APIs and OVHcloud’s AI Endpoints in France, both offering EU hosting with zero data retention and no training on customer prompts, the IONOS AI Model Hub and STACKIT in Germany, Nebius in the Netherlands, and Mistral’s La Plateforme, with aggregation services such as EUrouter routing traffic across EU-compliant endpoints (EUrouter, 2026; Infrabase, 2026). Their sovereignty, however, is layered and partial. The most capable models they serve are predominantly American or Chinese open-weight releases, the accelerators are Nvidia’s, and much of the underlying data centre capacity involves non-European operators or investors (Renda and Kyosovska, 2025). An EU inference provider can determine where the data ghost is processed, but it cannot yet ensure that the model performing that processing is itself European. This is precisely the gap the giga factory program is meant to close.

Read together, these instruments sketch a division of labor. The AI Continent Action Plan addresses sovereignty as capacity. Article 53 addresses it as provenance, while Opinion 28/2024 addresses it as residue. None yet delivers what the data ghost concept demands: an enforceable answer to the question of what happens when deletion is technically, legally, or institutionally incomplete.

Learning Platforms and Student – Faculty Ghosts

For education, the stakes are concrete because learning platforms (LMS) and AI tutoring systems do not capture only student data in the narrow administrative sense. They generate student ghosts through attendance logs, click paths, time-on-task, quiz attempts, error histories, revision patterns, writing samples, voice recordings, proctoring images, engagement scores, and predictions about risk, ability, persistence, or future performance. LMS’ also generate faculty ghosts by holding lecture recordings, slides, feedback comments, grading patterns, rubrics, course designs, LMS activity, response times, advising notes, student evaluations, and analytics about teaching effectiveness can all become durable traces of academic labor. These traces may outlive enrollment, employment, consent, and even the institutions that collected them.

A European sovereignty agenda worthy of the name should therefore be judged not only by the scale of its computational infrastructure, but by its capacity to answer a more immediate question: whether students and faculty can know where their traces have gone, how they are being recombined, and whether European law can follow the data ghost into the weights of the model.

 

References

European Commission. (2025a). The AI Continent Action Plan. https://commission. europa. eu/topics/competitiveness/ai-continent_en

European Commission. (2025b). Template for general-purpose AI model providers to summarise their training content, 24 July 2025. https://digital-strategy. ec. europa. eu/en/faqs/template-general-purpose-ai-model-providers-summarise-their-training-content

European Data Protection Board. (2024). Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models, 17 December 2024. https://www. edpb. europa. eu/system/files/2024-12/edpb_opinion_202428_ai-models_en. pdf

European Parliament and Council. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), Article 53. https://eur-lex. europa. eu/eli/reg/2024/1689/oj

European Parliament and Council. (2016). Regulation (EU) 2016/679 (General Data Protection Regulation), Article 17. https://eur-lex. europa. eu/eli/reg/2016/679/oj

EUrouter. (2026). AI providers. https://www. eurouter. ai/providers

Infrabase. (2026). AI infrastructure in Europe: EU-hosted tools and providers. https://infrabase. ai/european

Kyosovska, N., & Renda, A. (2025). Sanctuaries or cathedrals? CEPS In-Depth Analysis. https://cdn. ceps. eu/2025/11/251027-Sanctuaries-or-Cathedrals. pdf

author avatar
George Ellis
Share This Article