Master thesis: Erased or Suppressed? Auditing LLM Unlearning
- Hiring from
- Sweden
- Work type
- Hybrid
- Posted
- Oct 1, 2026
As Sweden's national center for applied AI, we're on a mission to accelerate the use of AI to benefit our society, our competitiveness, and everyone living in Sweden. We drive impactful initiatives in areas such as healthcare, energy, and public services while pushing the boundaries of AI research in fields such as natural language processing, machine learning and AI security. Join us in harnessing the untapped value of AI to drive innovation and create sustainable value for Sweden.
We are now looking for a master thesis student to join our team.
Introduction
When large language models (LLMs) memorize private, copyrighted, or sensitive data, developers use approximate machine unlearning to remove targeted data influence without costly retraining the model from its full training corpus. Current evaluations often measure success behaviorally, treating information as forgotten if a model fails to produce the target answer under direct questioning [8-10]. However, surface refusal is deceptive: deleted facts routinely persist in intermediate hidden states or resurface under adversarial prompting, steering, and fine-tuning [2-5]. Without empirical stress-testing, unlearned models deployed in public risk regulatory exposure and data leakage. This thesis investigates whether machine unlearning genuinely erases target information or merely suppresses its output expression.
Project Background and Problem Statement
AI Sweden is leading the development of LeakPro, an open-source privacy auditing tool designed to assess information leakage risks in machine learning models [6]. This initiative, undertaken in collaboration with RISE, Sahlgrenska, Recorded Future, Region Västmanland, Halland, AstraZeneca, Syndata, and Scaleout, aims to evaluate the risk of sensitive information disclosure when models trained on confidential data are made publicly available
Existing benchmarks such as TOFU approximate the counterfactual with a single retained checkpoint and score behaviour at rest [1]. One reference model cannot separate genuine unlearning effects from ordinary seed-to-seed training variance, and at-rest scoring says nothing about stability under layer probing, quantization, or fine-tuning [2-6]. This thesis therefore asks: Can a multi-seed counterfactual audit detect residual target knowledge that TOFU-style behavioral evaluation reports as successfully forgotten, and for which classes of unlearning method does this gap between genuine erasure and behavioral suppression appear?
Outline
The goal of this project is to rigorously evaluate LLM machine unlearning against counterfactual baselines. The objectives are as follows:
1. Literature study of LLM unlearning and auditing: Summarize (i) approximate unlearning methods, (ii) evaluating unlearning though model outputs, internal layer memory, and relearning speed (iii) sequential unlearning dynamics, and (iv) suitable datasets and benchmark architectures [1, 5, 7].
2. Implementation of a counterfactual audit benchmark: Train a multi-seed counterfactual reference distribution on target-omitted data. Evaluate unlearned models across behavioral leakage, layer-level representation probes, and knowledge recovery audits [2-5].
3. Longitudinal and sequential unlearning analysis: Audit model behavior across repeated independent, semantically clustered, and overlapping deletion requests to construct a threat-model-specific erasure-suppression profile [7].
If time permits and the student has interest, there is also an opportunity to contribute to the open-source platform LeakPro currently under development [6] by integrating the unlearning audit outputs and taxonomy into LeakPro's privacy framework.
Who we’re looking for
We are seeking a curious, independent, and self-driven MSc student who wants to work at the absolute frontier of AI research.
Ongoing Master’s studies in Computer Science, Data Science, Engineering Physics, Complex Adaptive Systems, Machine Learning, or a related field.
Comfortable with Python and deep learning, as well as with the reality that an experiment might yield unexpected results.
At AI Sweden, we are committed to building diverse and inclusive teams. Some positions may be subject to export control regulations, which means that specific requirements may apply.
Why should you do your thesis with AI Sweden?
Doing your thesis at AI Sweden means working alongside leading AI scientists and change leaders. AI Sweden is Sweden’s National Center for AI, we drive research questions that have both a long shelf-life and are widely applicable to Swedish industry and the public sector. We aim for publications at the most competitive venues and celebrate a culture of research excellence.
As an organization, we’re uniquely positioned at the sweet spot of governmental influence and startup agility. Small enough to stay adaptive and have fun but backed by and in close contact with both the government, academia and private and public sector.
Practical details
This Master’s thesis at AI Sweden offers a hybrid setup in Gothenburg.
Application Deadline: 2026-10-25 (rolling selection – position may be filled earlier)
Start Date: January 2027
Contact
If you have any questions or thoughts, don’t hesitate to contact:
Fazeleh Hoseini, Research Scientist
Muhaddisa Ali, Research Scientist
AI Sweden does not accept unsolicited support and kindly ask not to be contacted by any advertisement agents, recruitment agencies or manning companies.