Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Intellias logo

Agent Evaluation Engineer

Intellias
Posted 3 hours ago
🇪🇬Egypt🏠Remote📁Engineering & Development
Is this job info correct?

Agent Evaluation Engineer

Location: Remote / Global

Engagement: Staff Augmentation through Intellias

Department: Core Architecture / Digital Engineering Platform

Degree: Bachelor’s Degree

Experience: 4+ years


About the Role

We are looking for an Agent Evaluation Engineer to design and build enterprise-grade evaluation frameworks for agentic AI systems.

In this role, you will develop automated evaluation pipelines, deployment quality gates, and reliability measurement frameworks that ensure safe, predictable, and production-ready agent behavior.

You will work closely with AI, platform, and engineering teams to establish robust testing methodologies covering reasoning quality, tool selection and usage, multi-turn interactions, workflow execution, and production feedback integration.

The ideal candidate has strong experience building automated evaluation or testing frameworks for ML, LLM, or agentic AI systems, with hands-on knowledge of LangGraph or comparable agent orchestration frameworks.


Project Overview

Our customer is a multinational corporation with more than a century of history, operating in 180+ countries and serving more than 1 billion consumers worldwide.

The organization is undertaking a major transformation focused on introducing a new generation of Reduced-Risk Products (RRPs) and developing an innovative digital ecosystem around IoT, eCommerce, digital marketing, and AI.

Its IT platform hosts 700+ applications, creating significant requirements for scalable, secure, reliable, and standardized engineering capabilities.


Intellias supports the engineering of a comprehensive software ecosystem for a game-changing IoT product at the intersection of innovative consumer experiences and cutting-edge technology.

As an Agent Evaluation Engineer, you will join the Core Architecture Team and contribute to the architecture and implementation of best practices across the Digital Engineering Enterprise Platform.

The platform provides engineering teams with reusable services, technologies, and practices that accelerate software development and operations while ensuring compliance, reliability, and engineering quality.


Key Responsibilities

Agent Evaluation Frameworks

  • Design and implement automated evaluation frameworks for LangGraph-based agent workflows and orchestration pipelines.
  • Build reusable test harnesses for LangGraph graphs, nodes, state transitions, and agent execution paths.
  • Develop build-time evaluation suites covering:
  • Tool selection and usage accuracy
  • Agent trajectory accuracy
  • Reasoning quality
  • Final output quality
  • Establish evaluation metrics, benchmarks, and acceptance criteria for agent releases.
  • Design scalable evaluation architectures that can be reused across multiple agent applications.

Multi-Layer Evaluation

  • Design evaluation frameworks combining deterministic graders and LLM-as-judge graders.
  • Implement a three-layer evaluation model covering:
  1. Tool selection and trajectory accuracy
  2. Reasoning quality
  3. Final output quality
  • Develop scoring methodologies that balance deterministic validation with semantic and qualitative evaluation.
  • Ensure evaluation results are reproducible, measurable, and actionable.

Multi-Turn & Conversation Testing

  • Design automated simulations for multi-turn agent conversations.
  • Evaluate context retention, memory utilization, and consistency across multiple interactions.
  • Develop test scenarios covering complex workflows, state transitions, tool calls, and agent decisions.
  • Identify regressions in agent behavior across conversation turns.
  • Build repeatable datasets and test cases for evaluating agent performance over time.

Reliability & Statistical Evaluation

  • Define and implement reliability measurement methodologies for agentic systems.
  • Implement metrics such as:
  • Pass@k – probability of achieving success at least once across multiple trials.
  • Pass^k – probability of achieving consistent success across consecutive trials.
  • Design multi-trial evaluation strategies to measure both success rates and behavioral consistency.
  • Establish reliability thresholds and release criteria based on evaluation results.
  • Analyze evaluation trends and identify areas for agent improvement.

CI/CD Quality Gates

  • Design and implement CI/CD deployment gates based on agent evaluation metrics.
  • Automatically block releases when evaluation results fall below defined quality thresholds.
  • Integrate evaluation frameworks into build and deployment pipelines.
  • Implement validation processes covering:
  • Staging environments
  • Shadow-mode traffic comparisons
  • A/B testing and controlled rollouts
  • Ensure evaluation results become an integral part of the software delivery lifecycle.

AWS AgentCore Evaluations

  • Integrate AWS AgentCore Evaluations into continuous testing and quality assurance workflows.
  • Support both on-demand and online evaluation modes.
  • Leverage production evaluation results to improve build-time evaluation suites.
  • Where applicable, develop and integrate custom evaluators.
  • Establish feedback loops between production agent behavior and pre-production evaluation.

Production Feedback & Regression Testing

  • Transform production incidents, failures, and unexpected agent behavior into automated regression test cases.
  • Build feedback mechanisms that continuously improve evaluation datasets and test coverage.
  • Analyze production behavior to identify recurring reliability and quality issues.
  • Connect observability and production feedback with the evaluation framework.
  • Ensure critical production failures are represented in future release validation.

Collaboration & Documentation

  • Collaborate with AI engineers, platform teams, DevOps engineers, and product stakeholders to improve agent reliability and performance.
  • Define quality standards, acceptance criteria, and evaluation methodologies.
  • Participate in architecture discussions and technical design reviews.
  • Produce documentation covering evaluation strategies, scoring methodologies, test frameworks, deployment gates, and quality standards.
  • Promote best practices for safe and reliable agent development across the organization.


Required Qualifications & Experience

  • Bachelor’s degree in computer science, Software Engineering, Artificial Intelligence, Data Science, or a related field.
  • 4+ years of experience building automated testing or evaluation frameworks for ML, LLM, or agentic AI systems.
  • Hands-on experience designing multi-layer evaluation frameworks.
  • Experience combining deterministic graders and LLM-as-judge evaluation techniques.
  • Experience implementing automated CI/CD quality gates based on evaluation metrics and thresholds.
  • Hands-on experience with LangGraph or a comparable agent orchestration framework.
  • Strong understanding of agent evaluation concepts, including tool usage, reasoning, trajectories, and output quality.
  • Experience designing automated test harnesses and evaluation pipelines.
  • Experience with multi-turn conversation testing and context-retention evaluation.
  • Understanding of reliability measurement and multi-trial evaluation methodologies.
  • Strong understanding of software testing, automation, and continuous integration practices.
  • Strong analytical and problem-solving skills.
  • Strong English communication and documentation skills.


Nice-to-Have Qualifications

  • Hands-on experience with AWS AgentCore Evaluations, including CreateEvaluation and custom evaluators.
  • Experience with shadow-mode or canary deployments for ML/AI systems.
  • Experience with AWS Bedrock or other enterprise AI platforms.
  • Experience building LLM evaluation datasets and benchmark suites.
  • Experience with production observability and AI monitoring.
  • Experience transforming production incidents into automated regression tests.
  • Experience with AI safety, reliability, or responsible AI evaluation.
  • Familiarity with statistical evaluation methodologies for probabilistic AI systems.


Key Technical Skills

Agent Evaluation | LLM Evaluation | AI Testing | LangGraph | Agentic AI | LLM-as-Judge | Deterministic Graders | Evaluation Frameworks | Test Automation | CI/CD Quality Gates | Pass@k | Pass^k | Multi-Turn Testing | Tool Selection Evaluation | Agent Trajectory Evaluation | AWS AgentCore Evaluations | Regression Testing | Shadow Testing | Canary Deployment | Production Feedback Loops

Why This Position?

  • Shape the quality standards for enterprise AI: Build the evaluation foundation that determines whether AI agents are safe and reliable enough for production.
  • Work on cutting-edge agentic AI: Develop evaluation methodologies for multi-step, tool-using, stateful AI agents.
  • Real production impact: Connect build-time testing with real-world agent behavior and production feedback.
  • Architecture influence: Join the Core Architecture Team and help establish enterprise standards for AI quality and reliability.
  • Global scale: Support technology platforms spanning 700+ applications across 180+ countries.
  • Automation-first environment: Build evaluation systems that become an integral part of the software delivery lifecycle.
  • Continuous improvement: Turn production failures and unexpected agent behavior into actionable regression tests.
  • Intellias partnership: Join through Intellias and contribute to a major global technology transformation.

Education

Bachelor’s degree in computer science, Software Engineering, Artificial Intelligence, Data Science, or a related technical field.

Similar jobs

Similar jobs

Careerflow logo

Data Annotation Specialist - Software Engineer (OpenClaw Trajectory Evaluation)

Careerflow

🌍Bangladesh, Brazil, Egypt, Ghana, India, Kenya, Mexico, Nigeria, Pakistan, TurkeyMay 27, 2026, 10:54 PM UTC
ArkLeap Technologies logo

Oracle CX / CPQ Consultant

ArkLeap Technologies

🇪🇬Egypt3 hours ago
N|

Backend Developer (Node.js)

NUGTTAH | نقطة

🇪🇬Egypt3 hours ago
Intellias logo

Senior Python Backend Engineer – A2A & AWS AgentCore

Intellias

🇪🇬Egypt3 hours ago
Intellias logo

Strong Middle Python Engineer (MCP)

Intellias

🇪🇬Egypt3 hours ago
Slihrms logo

BIM Engineer II _ (Infra. Wet Utilities)

Slihrms

🇪🇬Egypt13 hours ago