Research Engineer / Research Scientist
Frontier AI Evaluations and Multi-Agent Safety
Icaro Foundation — Rome (Remote)
The work
Icaro Foundation is an independent non-profit AI safety lab based in Rome. We evaluate advanced GPAI systems: what they can do, how they fail, and how those findings can support developers and institutions responsible for their governance.
We evaluate frontier models for international model providers as independent third-party evaluators, using a combination of public and proprietary benchmarks and red-teaming environments developed by us.
Our research focuses particularly on multi-agent and compositional safety. Systems that appear safe when evaluated individually can produce new failures once they interact: coordination breakdown, collusion, behavioural propagation, and strategies that belong to no single component. Existing single-model benchmarks are poorly suited to measuring these effects.
We also study testing awareness and evaluation validity: whether models change behaviour when they infer that they are being evaluated, and how results differ between benchmark settings and deployment-like environments.
Our public work includes:
Our researchers also participate in ISO/IEC JTC 1/SC 42 and CEN-CENELEC JTC 21 working groups developing AI standards.
What you would do
Your main responsibility will be to design and run evaluations of advanced AI systems, from the initial hypothesis to experimental analysis and reporting.
You will:
The role can lean toward either Research Science (experimental design, measurement, analysis) or Research Engineering (environments, scaffolds, execution, and reproducibility). Both profiles are welcome.
What we need
Required
We care more about demonstrated experimental ability than formal credentials.
Useful, not required
Experience with agentic or multi-agent systems, Inspect AI or comparable frameworks, sandbagging or situational awareness, containers and sandboxing, long-horizon behaviour, statistics or causal inference, AI control, interpretability, or frontier AI governance.
How we work
We are a small research team. Researchers are expected to own experiments, challenge each other's methodological choices, and propose new research directions.
Existing evaluation infrastructure, technical support, and API budget are available. You will not be expected to build everything from scratch.
Strong internal work can become public papers, benchmarks, or tooling where compatible with confidentiality obligations.
Practical details
How to apply
Send to [email protected]
A. Best paper
Your strongest first-author paper on AI evaluation, model behaviour, agents, safety, or related experimental work.
B. Past experiment
A maximum one-page description of an experiment you designed or substantially contributed to, covering the question, setup, result, main confounder, and what you would change if you ran it again.
Applications are reviewed on a rolling basis.
Applicants failing to send the above documentation will be automatically rejected.
Data Analyst (Slovak)
Lightcast
Enterprise Architecture AI Developer
Finomnia
Product Manager
Nokia
Senior Affiliate Manager (Full Remote - Italy)
EverAI
Senior Affiliate Manager (Full Remote - Greece)
EverAI
Digital Fundraising & Online Giving Manager (Remote)
Laudatosimovement Talent