AS

Masters Thesis: Input-Level Causal Attribution in LLM Agents

Hiring from
Sweden
Work type
Hybrid
Posted
Oct 1, 2026
Is this job info correct?

As Sweden's national center for applied AI, we're on a mission to accelerate the use of AI to benefit our society, our competitiveness, and everyone living in Sweden. We drive impactful initiatives in areas such as healthcare, energy, and public services while pushing the boundaries of AI research in fields such as natural language processing, machine learning and AI security. Join us in harnessing the untapped value of AI to drive innovation and create sustainable value for Sweden.

We are now looking for a master thesis student to join our team.

Introduction

LLM agents increasingly act on retrieved documents, tool outputs, and memory from earlier sessions. When such an agent takes an incorrect or harmful action, it is often unclear which input led to it. Observability platforms have begun to offer automated root-cause suggestions, but their reliability is hard to assess, since real incidents rarely come with a known cause.

Existing failure-attribution benchmarks, such as Who&When dataset (Zhang et al., ICML 2025) and TraceElephant (Chen et al., 2026), identify the responsible agent or step rather than the input content, and step-level accuracy remains low. Context-attribution methods such as AttnTrace (arXiv:2508.03793) are mostly evaluated on single injection attacks. This thesis builds a benchmark of executed agent traces with planted, verified causes and uses it to evaluate existing attribution methods.

Background and Problem Statement

AI Sweden leads LeakPro (github.com/aidotse/LeakPro), an open-source tool for assessing data leakage from machine learning models and synthetic data. The LeakPro II consortium includes RISE, Sahlgrenska University Hospital, Recorded Future, Region Västmanland, Region Halland, AstraZeneca, Syndata and Scaleout.

Agents open a new leakage path: retrieved content, tool output or stale memory can steer an agent into disclosing data it should not. Benchmarks such as InjecAgent and AgentDojo measure whether such attacks succeed. An audit also needs the second question answered: after a disclosure, can its cause be traced?

Research question: Given a recorded agent trace and an action, how accurately can current attribution methods identify the input segment or segments that caused it, and how does accuracy depend on cause type, number of interacting causes, trace length and sampling non-determinism?

Outline

The thesis has three parts.

  1. Literature study. Review how current methods find which step or input caused an agent to fail or misbehave.

  2. Build a test set with known answers. Build a simple agent that searches documents, calls tools, and remembers past sessions. Its document store contains fake sensitive records. Plant an input that makes the agent leak one of them: a hidden instruction in a document, a false tool result, or an outdated memory. Confirm the planted input really causes the leak by running the agent many times with and without it. Because the cause is planted, the correct answer is known for every case. Include cases with two causes working together, and cases with no planted cause at all.

  3. Test existing methods. Give each method the recorded agent run and ask which input caused the leak. Measure how often it finds the planted input, how much computation it needs, and how results change with longer runs and more random model behaviour.

Who we’re looking for

  • We are seeking a curious, independent, and self-driven MSc student who wants to work at the absolute frontier of AI research.

  • Ongoing Master’s studies in Computer Science, Data Science, Engineering Physics, Complex Adaptive Systems, Machine Learning, or a related field.

  • Comfortable with Python and deep learning, as well as with the reality that an experiment might yield unexpected results.

At AI Sweden, we are committed to building diverse and inclusive teams. Some positions may be subject to export control regulations, which means that specific requirements may apply.

Why should you do your thesis with AI Sweden?

Doing your thesis at AI Sweden means working alongside leading AI scientists and change leaders. AI Sweden is Sweden’s National Center for AI, we drive research questions that have both a long shelf-life and are widely applicable to Swedish industry and the public sector. We aim for publications at the most competitive venues and celebrate a culture of research excellence.

As an organization, we’re uniquely positioned at the sweet spot of governmental influence and startup agility. Small enough to stay adaptive and have fun but backed by and in close contact with both the government, academia and private and public sector.

Practical details
This Master’s thesis at AI Sweden offers a hybrid setup in Gothenburg.

Application Deadline: 2026-10-25 (rolling selection – position may be filled earlier).
Start Date: January 2027

Contact

If you have any questions or thoughts, don’t hesitate to contact:

Fazeleh Hoseini, Research Scientist

AI Sweden does not accept unsolicited support and kindly ask not to be contacted by any advertisement agents, recruitment agencies or manning companies.

Similar jobs

Apply for this job