Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced QA and Test Engineers to ensure every benchmark is reliable, reproducible, and accurately measures real AI capabilities.
In this role, you will review complex, multi-step evaluation tasks, validate their correctness, identify edge cases, and strengthen testing methodologies before benchmarks are deployed. You'll collaborate closely with AI researchers and task authors to improve evaluation quality, eliminate ambiguity, and ensure benchmark integrity.
This is a fully remote, full-time engagement requiring approximately 35 hours per week.
Requirements
Key Responsibilities
Design comprehensive test cases that validate evaluation tasks, including complex edge cases and unexpected scenarios
Review benchmark tasks and reference solutions to identify ambiguity, inconsistencies, missing requirements, and grading gaps
Debug task environments and Python-based evaluation scripts to ensure reliable execution and accurate results
Develop repeatable quality assurance processes, validation checklists, and testing frameworks for benchmark creation
Identify potential shortcuts, exploits, or weaknesses that could compromise evaluation accuracy or benchmark integrity
Collaborate with AI researchers, engineers, and task authors to improve task quality, reproducibility, and technical rigor
Required Qualifications
Master's degree, PhD, or equivalent practical experience in a STEM discipline involving software engineering, research, or advanced technical problem solving
Minimum 1 year of professional experience in Quality Assurance, Test Engineering, Software Engineering, Research Engineering, or a related technical field with strong quality ownership
Proven experience designing test cases, validating complex software systems, and debugging end-to-end workflows
Strong proficiency in Python and Git, with the ability to troubleshoot unfamiliar codebases and technical environments
Excellent analytical thinking, problem-solving skills, and exceptional attention to detail
Experience documenting bugs, test strategies, and technical findings with clear written communication
Experience evaluating AI systems, machine learning models, or AI-generated outputs is preferred
Ability to work independently while managing multiple complex tasks with minimal supervision
Ability to commit approximately 35 hours per week on a consistent basis
Preferred Qualifications
Experience with AI evaluation, benchmark development, or quality assurance for machine learning systems
Background in automation testing, validation frameworks, or software quality engineering
Familiarity with large language models, AI agent workflows, or evaluation pipelines
Experience creating repeatable QA processes for research or engineering projects
Why Join
Play a key role in improving the quality and reliability of next-generation AI evaluation benchmarks
Collaborate with leading AI researchers on cutting-edge evaluation methodologies
Help ensure AI systems are tested against rigorous, real-world scenarios
Apply your software testing and quality engineering expertise to advance AI reliability
Enjoy the flexibility of a fully remote engagement while contributing to impactful AI research
Equal Opportunity
We are committed to creating an inclusive workplace where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.
Contract & Engagement Details
Independent contractor engagement
Fully remote with flexible working hours
Expected commitment of approximately 35 hours per week
Project duration may be extended, shortened, or concluded based on project requirements and individual performance
Work does not require access to confidential or proprietary information from any current or former employer
Payments are issued weekly based on approved work completed
At this time, we are unable to support H1-B or STEM OPT candidates