Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact mahmoud@relomote.com · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Clera logo

Data Scientist — Agent Evaluations & Quality

Clera
Posted 1 hour ago
🇺🇸United States🏠Remote📁Data & Analytics
Is this job info correct?

About the Role This is an applied data science role focused on measuring, understanding, and improving the quality of AI agents that handle real-world tasks — scheduling, email, browser automation, business software, and more. You'll sit at the intersection of evaluation design, statistics, and production systems, building the feedback loops that drive engineering and product decisions. The work matters because agent quality is hard to measure and easy to get wrong. What You'll Do Architect and maintain automated evaluation pipelines that measure agent quality across capabilities and product surfaces. Translate complex, multi-step agent behaviors into explicit success criteria — including pass, partial-pass, and failure definitions. Build gold datasets and regression suites covering common workflows, edge cases, and adversarial scenarios. Define and track metrics spanning task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability. Design deterministic and model-based graders; calibrate LLM-as-a-judge systems and measure false positives, false negatives, variance, and grader agreement. Analyze traces, tool calls, and production outcomes to identify root causes and build a useful failure taxonomy. Run rigorous offline experiments and leverage production evidence to compare models, prompts, and capability implementations. Build dashboards and reports that make evaluation results clear and actionable for engineering, product, and leadership. Partner with engineers to recommend improvements and verify fixes raise quality without unacceptable regressions. What We're Looking For 5+ years in data science, machine learning, or analytics roles, with a focus on evaluation systems, metrics frameworks, or quality measurement for production systems. Demonstrated experience designing and implementing evaluation frameworks and grading systems for ML or AI products in production. Production-quality Python and SQL skills; ability to build automated pipelines and conduct analysis at scale. Strong evaluation methodology chops: success criteria definition, dataset construction, metric selection, and spotting misleading benchmarks. Solid statistical and experimental design knowledge — sampling, variance, uncertainty quantification, bias, confounding, and significance testing for non-deterministic systems. Experience developing ground-truth data: labeling guidelines, annotation QC, ambiguity resolution, and dataset maintenance. Working knowledge of LLM behavior, tool use, retrieval systems, multi-step execution, and practical failure modes. Ability to connect quantitative patterns to individual traces and identify failure origins across model, prompt, context, tools, and application logic. Clear communicator who can convey evaluation results, methodology, and trade-offs to both technical and non-technical stakeholders. Experience with agentic systems, multi-step task evaluation, or consumer-facing production ML is a strong plus. Location On-site in Palo Alto, California, United States. Visa sponsorship is not available.

Similar jobs

Similar jobs

Bioscope AI logo

Data Scientist (AI Quality & Evaluation)

Bioscope AI

🇺🇸United States1 weeks ago
Bme Strategies logo

(Open Application) Consultant, Monitoring, Evaluation & Quality Improvement

Bme Strategies

🇺🇸United States1 weeks ago
PR

Vice President, Program Evaluation and Quality Assurance

Projectrenewal

🇺🇸United States1 weeks ago
Alpinephysicians logo

Clinical Quality and Program Evaluation Specialist

Alpinephysicians

🇺🇸United States2 weeks ago
Harvey logo

Senior Product Operations Manager, Evaluation Quality

Harvey

🇺🇸United States4 weeks ago
Bme Strategies logo

Consultant, Monitoring, Evaluation & Quality Improvement (Open Application)

Bme Strategies

🇺🇸United StatesJul 15, 2026, 4:10 AM UTC