Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Clera logo

Data Scientist — Agent Evaluations & Quality

Clera
Posted 6 hours ago
🇺🇸United States🏠Remote📁Data & Analytics
Is this job info correct?

About the Role This is an applied data science role focused on measuring, understanding, and improving the quality of AI agents that handle real-world tasks — email, calendar, browser, and business software. You'll sit at the intersection of evaluation design, statistics, and production engineering, building the feedback loops that directly guide how the product and engineering teams make decisions. Getting agent quality measurement right is core to how this product improves. What You'll Do Architect and maintain automated evaluation pipelines that measure agent quality across capabilities and product surfaces. Translate agent capabilities into explicit success criteria, including pass, partial-pass, and failure definitions for complex multi-step tasks. Build representative gold datasets and regression suites covering common workflows, edge cases, long-tail behavior, and adversarial scenarios. Define and track metrics such as task success, tool-selection accuracy, instruction adherence, factual consistency, latency, cost, and reliability. Design deterministic and model-based graders; calibrate LLM-as-a-judge systems and measure false positives, false negatives, variance, and grader agreement. Analyze traces, tool calls, model outputs, and production outcomes to identify root causes and build a useful failure taxonomy. Compare models, prompts, tools, and capability implementations using rigorous offline experiments and production evidence. Build dashboards and release-quality signals that make evaluation results actionable for engineering, product, and leadership. Partner with capability engineers to verify that fixes improve quality without unacceptable regressions in cost, latency, or reliability. What We're Looking For 5+ years in data science, machine learning, or analytics roles delivering evaluation systems, metrics frameworks, or quality measurement for production systems. Demonstrated experience designing and implementing evaluation frameworks and grading systems for ML or AI systems in production. Production-quality Python and SQL; ability to build automated data pipelines and analysis code at scale. Strong evaluation methodology skills: success criteria definition, dataset construction, metric selection, and identifying misleading benchmarks. Statistical and experimental design knowledge including sampling, variance, uncertainty quantification, bias detection, and significance testing for non-deterministic systems. Experience with ground-truth data development: labeling guidelines, annotation quality control, ambiguity resolution, and dataset maintenance. Working knowledge of LLM behavior, tool use, retrieval, multi-step execution, and practical failure modes of language model systems. Ability to connect quantitative patterns to individual system traces and identify failure origins across model, prompt, context, tools, and application logic. Experience building dashboards and communicating evaluation results, methodology, and trade-offs to both technical and non-technical stakeholders. Familiarity with LLM-as-a-judge systems, agentic pipelines, or benchmarking platforms for AI is a strong plus. Location On-site in Palo Alto, CA. Visa sponsorship is not available for this role.

Similar jobs

Similar jobs

Bioscope AI logo

Data Scientist (AI Quality & Evaluation)

Bioscope AI

🇺🇸United States2 days ago
Bme Strategies logo

(Open Application) Consultant, Monitoring, Evaluation & Quality Improvement

Bme Strategies

🇺🇸United States4 days ago
PR

Vice President, Program Evaluation and Quality Assurance

Projectrenewal

🇺🇸United States5 days ago
Alpinephysicians logo

Clinical Quality and Program Evaluation Specialist

Alpinephysicians

🇺🇸United States6 days ago
Harvey logo

Senior Product Operations Manager, Evaluation Quality

Harvey

🇺🇸United States2 weeks ago
Bme Strategies logo

Consultant, Monitoring, Evaluation & Quality Improvement (Open Application)

Bme Strategies

🇺🇸United StatesJul 15, 2026, 4:10 AM UTC