AI Sweden logo

Masters Thesis: GRPO for Multilingual Math and Logic Reasoning

Hiring from
Sweden
Work type
Hybrid
Posted
Oct 2, 2026
Is this job info correct?

As Sweden's national center for applied AI, we're on a mission to accelerate the use of AI to benefit our society, our competitiveness, and everyone living in Sweden. We drive impactful initiatives in areas such as healthcare, energy, and public services while pushing the boundaries of AI research in fields such as natural language processing, machine learning and AI security. Join us in harnessing the untapped value of AI to drive innovation and create sustainable value for Sweden.

We are now looking for a master thesis student to join our team.

Introduction

Post-training approaches using Reinforcement Learning (RL) are critical for eliciting deep reasoning capabilities in LLMs. Group Relative Policy Optimization (GRPO) removes the need for a separate value model, drastically reducing compute overhead. However, the efficacy of GRPO on non-English reasoning tasks remains largely unexplored.

Project Background and Problem Statement

OpenEuroLLM is integrating the Dolci-Think-SFT datasets into its post-training pipeline. We need to validate whether GRPO can effectively utilize translated and native math/logic rollouts to enhance Prelude 9B's reasoning capabilities in Swedish and other EU languages. The thesis asks:

  • Does GRPO on translated Swedish math problems yield compute-optimal reasoning improvements compared to standard SFT?

Outline

The goal is to apply compute-efficient RL to improve Prelude 9B's reasoning capabilities.

  1. Literature study: Analyze RLHF, PPO, GRPO methodologies, and reasoning benchmarks (e.g., GSM8k, MATH).

  2. Implementation: Adapt the oellm-rlvr framework to perform GRPO on Prelude 9B using Swedish reasoning datasets (e.g., translated Dolci-Think).

  3. Evaluation: Benchmark the resulting model on multilingual reasoning tasks, explicitly measuring the trade-off between training TFLOPs/s, memory constraints, and final accuracy.

Who we’re looking for

We are seeking curious, self-driven MSc students eager to work at the frontier of open-weight European AI research (LLMs). You thrive on empirical discovery, design rigorous experiments, and let data challenge your assumptions.

  • Ongoing Master’s studies in Computer Science, Data Science, Machine Learning, Engineering Physics, or a related quantitative field.

  • Proficiency in Python and hands-on experience with modern deep learning frameworks (PyTorch, Hugging Face ecosystem).

  • Familiarity with LLM post-training alignment (e.g., SFT, DPO, RLHF/RLVR) or context-extension, alongside comfort running distributed GPU training in Linux/HPC environments.

At AI Sweden, we are committed to building diverse and inclusive teams. Some positions may be subject to export control regulations, which means that specific requirements may apply.

Why should you do your thesis with AI Sweden?

Doing your thesis at AI Sweden means working alongside leading AI scientists and change leaders. AI Sweden is Sweden’s National Center for AI, we drive research questions that have both a long shelf-life and are widely applicable to Swedish industry and the public sector. We aim for publications at the most competitive venues and celebrate a culture of research excellence.

As an organization, we’re uniquely positioned at the sweet spot of governmental influence and startup agility. Small enough to stay adaptive and have fun but backed by and in close contact with both the government, academia and private and public sector.

Practical details
Location: Hybrid (Gothenburg / Stockholm) or Remote.

Application Deadline: 2026-11-20 (rolling selection – position may be filled earlier).
Start Date: January 2027

Contact

If you have any questions or thoughts, don’t hesitate to contact:

Amaru Cuba Gyllensten, Senior Research Scientist

Birger Moëll, Senior Research Scientist

AI Sweden does not accept unsolicited support and kindly ask not to be contacted by any advertisement agents, recruitment agencies or manning companies.

References

[1] Shao, Z., et al., "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models," arXiv 2024.

[2] OpenEuroLLM Consortium, "oellm-rlvr Framework Documentation," 2026.

Similar jobs

Apply for this job