Role Overview:
We are hiring for one of our clients, seeking an AI Benchmark Engineer — Software Engineering to work on a full-time basis. The role involves designing and building high-quality multi-agent benchmark tasks grounded in real-world open-source code changes such as bug fixes, migrations, and refactors. You will evaluate how effectively AI agents understand large codebases, apply precise modifications, and produce correct, testable outputs.
Key Responsibilities:
• Design and build multi-agent benchmark tasks based on real-world open-source code changes, including bug fixes, migrations, and refactors.
• Work within the Harbor evaluation framework to run and validate tasks inside Docker environments.
• Write clear, precise task instructions specifying file paths, function signatures, expected behavior, and constraints.
• Develop Python-based verification scripts to validate correctness of agent-generated code changes.
• Decompose complex engineering problems across multiple specialized agents and refine tasks within containerized environments.
Required Skills & Qualifications:
• 5+ years of experience in Python and JavaScript development.
• Experience with AI coding benchmarks such as SWE-bench or Terminal-bench.
• Strong ability to read and navigate large open-source codebases, including frameworks like Django, Flask, FastAPI, or Node.js.
• Familiarity with Git workflows, including pull requests, diffs, cherry-picking, and working with specific commits.
• Experience with Docker, including writing Dockerfiles and managing builds.
More About the Opportunity:
This role offers a unique opportunity to work with a global leader in the Technology, Information and Internet industry, contributing to the advancement of AI systems that solve mission-critical priorities. The work directly supports the evaluation of cutting-edge AI agents in real-world software engineering scenarios.
Equal Opportunity Employer:
We hire based on skills and expertise. All qualified candidates are welcome regardless of background, experience, or prior employment history. Applications are reviewed solely on demonstrated technical ability and qualifications.
Apply Now!
AI/ML Engineer (Remote)
Jobs Ai
Software Engineer - AI Code Evaluation (Remote)
Jobs Ai
Lead AI Engineer
Emumba
AI Engineer
Codeninjapk
AI Engineer
Remotebase
Electrical Engineering AI Reviewer (Signal Processing)
Gramian Consulting Group