Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Clera logo

Research Engineer, Benchmarks

Clera
Posted 5 hours ago
🛂Visa sponsorship
🇺🇸United States
💰$150.0K–$250.0K📁Engineering & Development
Is this job info correct?

About the Role This is a high-ownership research engineering role on a small, technical team focused on designing and building benchmarks that evaluate frontier AI agents on realistic, domain-specific workflows. You'll work alongside researchers and engineers to ensure evaluations are rigorous, credible, and trusted by AI labs and enterprise customers. The quality of these benchmarks directly shapes how the world measures and improves AI agent performance. What You'll Do Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks. Partner with subject-matter experts to define realistic workflows and translate them into evaluation tasks and criteria. Build and operate reliable infrastructure to run models and agents against benchmark tasks at scale. Develop metrics and statistical analyses to measure benchmark difficulty, reliability, and failure modes. Validate that benchmark performance correlates with real-world evaluations and frontier lab expectations. Write clear technical documentation and benchmark reports for research and engineering audiences. What We're Looking For 2–4 years of experience in research engineering, ML engineering, or related roles — with a focus on AI benchmarks, evaluation infrastructure, or agent environments. Strong proficiency in Python, Docker, and Linux environments for building research or production infrastructure. Experience designing and running benchmarks or evaluation environments for AI agents or large language models. Experience developing metrics, statistical analyses, or validation studies to assess benchmark quality and real-world correlation. Ability to collaborate with domain experts to translate complex workflows into well-scoped evaluation tasks. Deep intuition for what makes a benchmark realistic, reliable, and practically useful. Strong attention to detail with a habit of catching subtle inconsistencies and edge cases. Comfort working independently in fast-moving, unstructured, early-stage environments. Excellent written communication skills; able to make technical results legible across time zones and teams. Nice to have: Published papers or technical writing on AI benchmarking, model evaluation, or failure modes; experience with RL training pipelines or data generation; background at frontier AI labs, research institutions, or on widely used public benchmark projects. Compensation & Benefits Salary range: $150,000 – $250,000 USD annually . Equity offered. Visa sponsorship is available. Location On-site in San Francisco, CA, USA . In-person presence is expected.

Similar jobs

Similar jobs

HU

Research Engineer, Benchmarks

hud

🌍United States, SingaporeJul 22, 2026, 6:12 PM UTC
HUD logo

Research Engineer, Benchmarks

HUD

🌍United States, SingaporeJul 14, 2026, 1:07 AM UTC
PH

Founding Engineer - LLM Infra & Platform

Pax Historia

🇺🇸United States3 hours ago
PH

Founding Engineer - Harness Optimization

Pax Historia

🇺🇸United States3 hours ago
CM

Founding Engineer

cmux

🇺🇸United States3 hours ago
RE

Hardware Engineer

Relari

🇺🇸United States3 hours ago