Principal Evals Engineer - AI & Agentic Systems
oryxsearch.ioPrincipal Evals Engineer - AI & Agentic Systems
About the Organisation
Our client is building one of the region’s most ambitious government AI programmes, developing the AI infrastructure, applications and platforms required to operate AI-native public services at scale.
The systems being built must operate reliably across Arabic and English, within strict requirements around data sovereignty, privacy, security and regulated infrastructure.
These are production systems supporting consequential workflows and senior decision-makers. Whether an AI system works cannot simply be a matter of opinion.
It has to be measurable.
This role owns how that measurement happens.
Why Join
- Solve difficult problems in applied AI. Determine whether non-deterministic AI systems actually work across real-world, bilingual and high-stakes environments.
- Own the quality bar. Define the evidence required before model changes, prompt updates, framework migrations or architectural changes reach production.
- Small, senior engineering teams. Work alongside experienced AI, product and forward-deployed engineers with significant technical autonomy.
- Build the evaluation practice. Establish the evaluation architecture, tooling and engineering standards that teams across the programme will use.
- High-impact environment. Your work will influence AI systems deployed across major public-sector organisations and used at significant scale.
- Relocation supported. Relocation assistance is available for successful candidates and their families.
The Role
We are hiring a Principal Evals Engineer to own how the organisation determines whether its AI systems actually work.
You will set the direction for evaluation across a portfolio spanning AI assistants, retrieval systems, agentic workflows, voice applications and document intelligence, while building the infrastructure that turns “it seems better” into measurable evidence.
You own measurement and release evidence, partnering closely with Principal-level engineers responsible for AI architecture, product engineering, platform development and forward deployment.
Evaluating AI systems is fundamentally different from conventional software testing. The same input can produce different outputs, correctness can be subjective, and failures such as hallucination, poor grounding, behavioural drift and prompt injection cannot be captured through conventional assertions alone.
At Principal level, you will be expected to have already tackled these problems in production environments.
This is primarily an individual contributor role. Your influence comes from the infrastructure you build, the quality bar you establish and the engineering decisions your evidence enables.
What You Own
Evaluation strategy and architecture.
Define what gets measured, at which layer, using which methodology and how results feed into product and release decisions. Establish the reference architecture engineering teams build against.
The shared evaluation platform.
Build evaluation harnesses, golden-set management, dataset versioning, automated grading, behavioural regression detection and reporting infrastructure.
Grading you can trust.
Own judge-model selection, rubric design and calibration against human labels — including understanding when automated grading cannot be trusted.
Retrieval, agents and multilingual evaluation.
Measure grounding, citation correctness, tool use, multi-step reasoning and failure recovery. Build dedicated Arabic evaluation datasets and ensure judges are properly calibrated rather than assuming English evaluation methods transfer directly.
Evidence behind engineering decisions.
Model swaps, prompt changes, framework migrations and infrastructure decisions should be supported by defensible evaluation evidence before changing production behaviour.
Production and adversarial evaluation.
Build online evaluation, sampling, human review, drift detection and alerting alongside adversarial testing for prompt injection, jailbreak resistance and data leakage.
Quality gates and escapes.
Integrate AI evaluation with conventional test automation and CI/CD, turning production quality escapes into evaluations capable of detecting the same failure before release.
Evaluation culture.
Enable engineering teams to run rigorous evaluations independently and ensure evaluation is designed into systems from the beginning rather than added before launch.
Who You Are
Staff or Principal-level track record.
Years matter less than ownership. You have built evaluation, testing or quality infrastructure used at meaningful scale.
Deep AI evaluation experience.
You understand output quality, retrieval, grounding and regression detection across non-deterministic systems — and how measurement translates into product improvement.
LLM-as-judge practitioner.
You understand rubric design, judge selection, calibration against human labels and the failure modes that can make automated grading misleading.
Strong Python engineer.
You can build production-quality evaluation harnesses, frameworks and shared libraries other engineers depend upon. Working knowledge of TypeScript or Java is valuable.
Statistically literate.
You understand sampling, confidence intervals, inter-rater agreement and statistical significance well enough to determine whether an observed improvement is genuine.
Strong CI/CD engineering.
You understand framework design across API, web and data surfaces, including test selection, parallelisation and flake management at scale.
Hands-on and current.
You still build the harnesses, run experiments and investigate failures yourself, using modern AI-assisted development tools as part of your workflow.
Strongly Preferred
- Experience with LLM evaluation platforms and frameworks such as Ragas, DeepEval, Promptfoo, Braintrust, LangSmith or Langfuse.
- Arabic AI evaluation, including golden datasets, dialect coverage, right-to-left validation and judge calibration.
- Adversarial and security testing for AI systems, including structured red-teaming.
- Voice and conversational AI evaluation.
- Data pipeline and AI-service performance testing.
- Human evaluation programme design.
- Experience delivering technology within government, financial services or another regulated/high-assurance environment.
Technology Environment
The environment includes:
Languages: Python, TypeScript and Java.
AI Evaluation: Custom evaluation harnesses, LLM-as-judge patterns, golden datasets, Langfuse, Ragas, DeepEval and Promptfoo.
AI & Agent Systems: LangChain, LangGraph, Microsoft Agent Framework, pgvector and Qdrant.
Testing: Playwright, Cypress, Appium, Detox, contract testing and schema validation.
Data Quality: Great Expectations and Soda.
Performance: k6, JMeter and Locust.
Infrastructure: Docker, Kubernetes and Azure.
CI/CD & Observability: GitHub Actions, GitLab CI, Grafana and behavioural regression dashboards.
AI Development: Modern coding agents and AI-native development workflows.
The technology stack reflects current engineering preferences rather than a rigid mandated standard.
What This Role Is Not
Not a traditional management role.
This is a hands-on Principal engineering position rather than a line-management role.
Not conventional QA leadership with AI added.
Test automation and performance testing are relevant, but the centre of gravity is evaluating non-deterministic AI systems.
Not a research-only role.
Success is not measured by papers, test cases or dashboards.
It is measured by whether engineering and product decisions are made using reliable evidence — and whether failures are identified before users encounter them.