Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
IF

AI Safety Researcher

Icaro Foundation
Posted 6 hours ago
🇮🇹Italy🏠Remote📁Data & Analytics
Is this job info correct?

Research Engineer / Research Scientist

Frontier AI Evaluations and Multi-Agent Safety

Icaro Foundation — Rome (Remote)


The work

Icaro Foundation is an independent non-profit AI safety lab based in Rome. We evaluate advanced GPAI systems: what they can do, how they fail, and how those findings can support developers and institutions responsible for their governance.

We evaluate frontier models for international model providers as independent third-party evaluators, using a combination of public and proprietary benchmarks and red-teaming environments developed by us.

Our research focuses particularly on multi-agent and compositional safety. Systems that appear safe when evaluated individually can produce new failures once they interact: coordination breakdown, collusion, behavioural propagation, and strategies that belong to no single component. Existing single-model benchmarks are poorly suited to measuring these effects.

We also study testing awareness and evaluation validity: whether models change behaviour when they infer that they are being evaluated, and how results differ between benchmark settings and deployment-like environments.

Our public work includes:

  • Boiling the Frog, on multi-turn agentic safety (arXiv:2605.22643);
  • the Adversarial Humanities Benchmark, on adversarial robustness across linguistic and conceptual registers (arXiv:2604.18487);
  • research on LLM-to-LLM risks, multi-agent collusion, and interaction-level safety (arXiv:2512.02682, arXiv:2601.11369, arXiv.org:2607.22188).

Our researchers also participate in ISO/IEC JTC 1/SC 42 and CEN-CENELEC JTC 21 working groups developing AI standards.


What you would do

Your main responsibility will be to design and run evaluations of advanced AI systems, from the initial hypothesis to experimental analysis and reporting.

You will:

  • turn open-ended safety questions into measurable evaluation protocols;
  • design and run experiments on frontier and open-weight models;
  • build environments involving agents, tools, memory, and multi-step tasks;
  • study coordination, collusion, behavioural propagation, and testing awareness;
  • use strong elicitation and control conditions;
  • distinguish genuine safety signals from confounders and experimental artefacts;
  • analyse model traces and write technical reports;
  • contribute to papers, benchmarks, tooling, and evaluation infrastructure.

The role can lean toward either Research Science (experimental design, measurement, analysis) or Research Engineering (environments, scaffolds, execution, and reproducibility). Both profiles are welcome.


What we need

Required

  • Master's degree or equivalent technical background.
  • Typically at least two years of substantive hands-on experience in AI research, engineering, or evaluation.
  • Strong Python and practical experience with LLMs.
  • Experience designing and running experiments, not only applying existing benchmarks.
  • Understanding of elicitation, controls, confounders, and measurement validity.
  • Evidence of serious experimental work: papers, repositories, benchmarks, datasets, or technical reports.
  • Strong technical writing.

We care more about demonstrated experimental ability than formal credentials.


Useful, not required

Experience with agentic or multi-agent systems, Inspect AI or comparable frameworks, sandbagging or situational awareness, containers and sandboxing, long-horizon behaviour, statistics or causal inference, AI control, interpretability, or frontier AI governance.

How we work

We are a small research team. Researchers are expected to own experiments, challenge each other's methodological choices, and propose new research directions.

Existing evaluation infrastructure, technical support, and API budget are available. You will not be expected to build everything from scratch.

Strong internal work can become public papers, benchmarks, or tooling where compatible with confidentiality obligations.


Practical details

  • Location: Fully remote.
  • Eligible locations: Europe and China.
  • Working hours: substantial overlap with European working hours is required.
  • China: we explicitly encourage researchers and engineers based in China to apply; candidates based there should be prepared to adapt their working hours accordingly.
  • Compensation: We compensate our contractors above the EU average hourly rate for comparable roles and seniority. We will share with candidates the specific range in the first round of the interview.
  • Working language: English.
  • Start date: November maximum.


How to apply

Send to [email protected]

  1. your CV;
  2. relevant links such as GitHub, Google Scholar, or personal website;
  3. one work sample, choosing either:


A. Best paper

Your strongest first-author paper on AI evaluation, model behaviour, agents, safety, or related experimental work.


B. Past experiment

A maximum one-page description of an experiment you designed or substantially contributed to, covering the question, setup, result, main confounder, and what you would change if you ran it again.

Applications are reviewed on a rolling basis.

Applicants failing to send the above documentation will be automatically rejected.

Similar jobs

Similar jobs

Lightcast logo

Data Analyst (Slovak)

Lightcast

🌍Czech Republic, Italy, Slovakia3 hours ago
FI

Enterprise Architecture AI Developer

Finomnia

🇮🇹Italy3 hours ago
Nokia logo

Product Manager

Nokia

🌍Italy, United States2 hours ago
EverAI logo

Senior Affiliate Manager (Full Remote - Italy)

EverAI

🇮🇹Italy3 hours ago
EverAI logo

Senior Affiliate Manager (Full Remote - Greece)

EverAI

🇮🇹Italy3 hours ago
Laudatosimovement Talent logo

Digital Fundraising & Online Giving Manager (Remote)

Laudatosimovement Talent

🇮🇹Italy3 hours ago