QA Engineer Come help define what “good” AI actually looks like - and build the systems that make quality measurable. We’re working with a fast-growing AI company looking for a QA Engineer to own how the quality of its AI products is measured, tested and improved. This isn’t traditional QA — you’ll work at the intersection of AI evaluation, software engineering and product quality, building the frameworks, datasets and automated tests that determine whether LLM-powered features are genuinely accurate, reliable and ready for production. You’ll work closely with engineering and product teams to identify where AI systems fail, understand why, and turn those findings into measurable improvements. Must Haves Strong software engineering or technical background, ideally with hands-on experience working with AI/ML systems Experience building testing, evaluation or quality frameworks for AI-powered products Strong Python or similar programming experience, with the ability to build evaluation tooling and automation Good understanding of LLMs and Generative AI, including how model outputs should be tested and evaluated Experience creating evaluation datasets, test cases and regression suites for complex AI features Ability to analyse model outputs, identify failure patterns and root causes, and turn findings into measurable improvements Strong understanding of accuracy, reliability and consistency when evaluating AI systems Excellent analytical and problem-solving skills, particularly when dealing with ambiguous or difficult-to-measure problems Comfortable working closely with Engineering and Product teams and communicating technical findings clearly Nice to Haves Experience evaluating RAG / retrieval systems, including the quality and relevance of retrieved information Experience testing AI agents, tool-calling or multi-step AI workflows Knowledge of different LLM evaluation methodologies, metrics and benchmarking approaches Experience implementing automated AI regression testing within development or CI/CD workflows Experience with human-in-the-loop evaluation alongside automated testing Knowledge of prompt evaluation and comparing changes across different models or prompts Experience building internal AI evaluation tooling, dashboards or monitoring Previous experience working with AI products where accuracy and reliability are particularly important Department Software Development Locations Multiple locations Remote status Fully Remote Applicant tracking system by Teamtailor