About the team Neo Research (新衡) is an independent AI safety research organization based in Singapore. We study frontier risks in increasingly capable AI systems, with a particular focus on the open-weight model ecosystem and the rapidly growing frontier-model ecosystem in Asia. Some of the world’s most capable open-weight models are now being developed in Asia and deployed globally. Yet, they remain poorly understood from a frontier-safety perspective. We want to understand how to ensure their safety, and what new risks become important as the frontier changes. Our current work focuses on misalignment, loss of control, and harmful manipulation. We study questions such as whether models pursue unintended objectives, conceal problematic behaviour, recognize and adapt to evaluations, evade oversight, or become less safe as they are given greater autonomy. Our goal is to produce rigorous empirical evidence about risks that are important but difficult to measure. We are looking for Research Engineers to build the experimental systems needed to evaluate increasingly capable models: realistic agent environments, model and tool integrations, long-horizon evaluation infrastructure, and the systems needed to make complex experiments reliable and reproducible. Why join Neo Research Evaluating increasingly capable models requires more than running benchmarks. It means building realistic environments, giving models tools and autonomy, running experiments that may unfold over hundreds of interactions, and capturing enough information to distinguish genuine model behaviour from quirks or failures of the evaluation itself. You will work closely with researchers while owning substantial parts of the experimental system,from agent scaffolds and tool integrations to model sampling, observability, trajectory analysis, and reproducibility. In this role, you will go beyond implementing evaluations and actively help shape evaluation methodology. Your work will contribute directly to published research shared with AI Safety Institutes, frontier labs, policymakers, and model developers. We have presented our work to most major Chinese model developers and run a joint evaluations project with an AI Safety Institute. Our mission statement describes our current research directions. In this role, you would: Design, implement, and run evaluations of misalignment and loss-of-control risks in frontier models. Build realistic agent environments, scaffolds, and tool integrations for studying behaviour under increasing autonomy. Develop infrastructure for long-horizon experiments involving many model calls, actions, tools, and environment states. Own experimental reliability and reproducibility across model access, sampling, configuration, environment management, logging, and analysis. Work with research scientists to turn open-ended questions into tractable experiments, and identify cases where an apparent model behaviour is actually an artefact of the evaluation. Build tools for inspecting and analysing large collections of model trajectories and comparing behaviour across models and experimental conditions. Contribute experimental methodology, infrastructure, and technical analysis directly to published research. About you Essentials Strong Python and general software-engineering skills. Experience working with language models, agentic systems, or model evaluations. The ability to turn an underspecified research question into a reliable experimental setup. Strong engineering practices, particularly around reproducibility, testing, observability, and data integrity. The ability to investigate unexpected results across both the model and the surrounding infrastructure. Clear technical writing and communication skills. Comfort working in an early-stage environment where requirements and research directions evolve. Helpful, but not required Experience with AI safety or dangerous-capability evaluations. Experience with evaluation frameworks such as Inspect. Experience building agent scaffolds or tool-use environments. Experience with distributed inference, large-scale API-based experimentation, Docker, Kubernetes, or related infrastructure. Experience analysing large collections of model trajectories or transcripts. Familiarity with frontier-model safety reports. Mandarin reading or writing ability. You don't need to have worked in AI safety specifically. We're also interested in strong ML and research engineers whose technical expertise and engineering judgment could transfer strongly to this work. Role logistics & benefits Location: This role can be based in Singapore or remotely. Our team is globally distributed, so remote team members should be comfortable maintaining some working-hour overlap with colleagues across regions. Compensation: Our compensation takes location, experience, and level into account, with indicative salary ranges of $120,000–$180,000+. We may offer above this range for exceptional candidates. Benefits: Competitive benefits and leave policies.