Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Mercor logo

Member of Technical Staff, Enterprise Evals Platform

Mercor
Posted 7 hours ago
📦Relocation support
🇺🇸United States
💰$220.0K–$425.0K📁Engineering & Development
Is this job info correct?

About Mercor Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents. Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices. About the Role Enterprise agents are complex systems, and they only pay off when their work is reliable and economically viable. Evaluation is how you get both: checking correctness is the obvious case, and routing is the subtler one, since choosing a model against cost, latency, and quality requires quality to be measurable at all. Knowing where the bar sits is the hard part. You decompose real work, take the standard from the practitioners who hold it, and encode it so an agent cannot shortcut it. You will apply what Mercor has learned building benchmarks with domain experts, and devise new methods, so that evals and rubrics keep improving and so do the agents measured against them. That work only scales with a platform behind it. This is a platform engineering role with good depth of understanding in evals. You will build the verifiers, the environments agents are measured in, and the grading infrastructure that runs at scale, abstracted across customers, domains, and tasks so that every run becomes evidence the next agent inherits instead of starting over. Read more about how we think about this: Agent Eval Systems Responsibilities Define golden sets: decompose real tasks and encode the expert quality bar. Build verifiers over agent trajectories and outputs, calibrated and hard to game. Build the eval platform that runs offline environments, task suites, and grading at scale. Run loss analysis over production trajectories and turn failure modes into regression tests. Run the optimization loop across models, prompts, skills, and harnesses. Own the rollout gates that decide whether an agent change ships. Partner with the Enterprise Platform team and the Applied AI engineers embedded with customers. What We're Looking For Professional, academic, or research experience in agent engineering and evaluation, including how agent runtimes and harnesses produce a trajectory and where it fails. Experience building evaluation suites for LLM or agent systems, and familiarity with how benchmarks such as terminal-bench, tau-bench, and APEX are constructed and where they get gamed. Judgment about task and rubric design: turning a fuzzy notion of quality into something measurable, with agent or model improvements to show for it. Strong software engineering fundamentals, and the ability to work independently on ambiguous, loosely specified problems. Bonus: experience with Harbor environments and RL environments. Why Mercor Impact: No agent reaches an enterprise customer without clearing the bar you set. Learning: Eval work spanning frontier model measurement and live enterprise deployments, on production trajectories few teams get to see. Growth: Research and systems in one role, with fast paths to owning the eval system end to end. Benefits Up to $15k relocation bonus $10K housing bonus (if you live within 0.5 miles of our office) $1.5K monthly stipend for meals Generous equity grant vested over 4 years Free Equinox membership $200 monthly laundry reimbursement $200 monthly personal wellness reimbursement Health, Dental, Vision insurance

Similar jobs

Similar jobs

Legion Technical Solutions Llc logo

System Safety Engineer 3

Legion Technical Solutions Llc

🇺🇸United States3 hours ago
Elanco logo

Sr. Associate, Quality Assurance Validation

Elanco

🇺🇸United States4 hours ago
Bridgestone logo

Electrical Engineer

Bridgestone

🇺🇸United States4 hours ago
CI

Cleared Cloud / DevOps Engineer - L4

Cleared IT Positions

🇺🇸United States4 hours ago
RideNow Powersports logo

C-Level Service Technician - RideNow North Fort Worth

RideNow Powersports

🇺🇸United States4 hours ago
Cat logo

Principal Software Engineer - Digital Twin and Simulation

Cat

🇺🇸United States4 hours ago