Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Clera logo

Research Engineer, Benchmarks

Clera
Posted 6 hours ago
🛂Visa sponsorship
🇺🇸United States
💰$150.0K–$250.0K📁Engineering & Development
Is this job info correct?

About the Role Join a small, highly technical team of researchers and engineers — including International Olympiad medalists and published AI researchers — at an early-stage startup building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. As a Research Engineer, Benchmarks , you'll own the design and implementation of evaluations that frontier labs and enterprise customers rely on to measure real-world agent performance. This is a critical, high-ownership role at the intersection of research rigor and engineering execution. The company operates in the AI/ML evaluation and reinforcement learning infrastructure space, providing a platform for building, running, and scaling RL environments and post-training datasets. The team is based in San Francisco, CA and works on-site. Visa sponsorship is available. What You'll Do Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks. Partner with subject-matter experts to define realistic workflows and tasks for domain-specific evaluations. Build reliable infrastructure to run models and agents against benchmark tasks at scale. Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes. Validate that benchmark performance correlates with real-world evaluations, customer needs, and frontier lab expectations. Write clear documentation and benchmark reports that make results legible and credible to technical audiences. What We're Looking For Required 2–4 years of experience in research engineering, ML engineering, or related roles — with a focus on building and delivering AI benchmarks, evaluation infrastructure, or agent environments. Demonstrated experience designing, implementing, and running benchmarks or evaluation environments for AI agents or large language models. Strong proficiency in Python, Docker, and Linux environments for building research or production infrastructure. Experience building and operating infrastructure to reliably run AI models or agents against benchmark or evaluation tasks at scale. Experience developing metrics, statistical analyses, or validation studies to assess benchmark difficulty, reliability, and real-world correlation. Experience collaborating with subject-matter experts to translate domain workflows into benchmark tasks and evaluation criteria. Experience analyzing workflows across diverse technical or business domains to inform task design. Strong technical writing skills — able to produce benchmark reports and documentation for research and engineering audiences. Nice to Have Published papers or technical blog posts on AI benchmarking, model evaluation, or model failure modes. Experience with reinforcement learning training pipelines, data generation, or RL agent evaluation. Background at frontier AI labs, research institutions, or involvement in widely used public benchmark projects. Traits We Value Deep curiosity about how workflows operate across varied domains. Sharp attention to detail — a habit of spotting subtle inconsistencies and edge cases in task design. Ability to reason from first principles about task design, scoring, and failure modes. Comfort thriving in unstructured problem spaces and working independently in a fast-paced, early-stage environment. Excellent communication skills for collaborating across time zones and with technical teams. Compensation & Benefits Salary: $150,000 – $250,000 USD annually, depending on experience. Early-stage equity participation. Visa sponsorship available. Location This is an on-site role based in San Francisco, CA, United States . Candidates must be willing and able to work from the office. Fully remote arrangements are not available for this position.

Similar jobs

Similar jobs

HU

Research Engineer, Benchmarks

hud

🌍United States, Singapore3 weeks ago
HUD logo

Research Engineer, Benchmarks

HUD

🌍United States, SingaporeJul 14, 2026, 1:07 AM UTC
Hntb logo

Section Manager - Overhead Contact Systems (OCS)

Hntb

🇺🇸United States6 hours ago
Hntb logo

Traffic Signals Project Engineer

Hntb

🇺🇸United States6 hours ago
Cat logo

Senior Software Engineer

Cat

🌍India, Slovakia, United States6 hours ago
GL

Software Engineer I (Onsite)

Globalhr

🇺🇸United States6 hours ago