The Role:
We're seeking a Senior AI QA Test Automation Engineer to take a leading technical role in AI quality and test automation. You will work primarily with our AWS Bedrock AgentCore-based CX Agent Suite, a multi-agent customer-support system covering intent routing, FAQs, deposit-status queries, and human handoff. You will design and build the evaluation and QA tooling required to validate AI agents reliably, while also contributing to technical standards and best practices across QA.
This is a Senior Individual Contributor (IC) role with a strong technical and strategic focus. You'll work closely with QA leadership, Data Science, and Engineering to architect scalable AI evaluation systems, establish technical standards, and drive the evolution of our AI evaluation framework. Our current evaluation harness is based on DeepEval and is evolving towards a broader agentic evaluation architecture. You will have the opportunity to evaluate and prototype emerging agent-orchestration approaches, including Strands Agents, and contribute to architectural decisions around their adoption.
Act as the primary technical enabler for QA, building scalable AI/ML frameworks, libraries, and tooling for the CX Agent Suite and beyond
Design and build AI evaluation pipelines that assess the CX Agent Suite's multi-agent responses (Intent Detector, FAQ Agent, Missing-Deposits Agent) for accuracy, relevance, tone, hallucination rate, safety/guardrail compliance, and task completion — extending or replacing the current DeepEval-based evaluation harness
Evaluate and prototype Strands Agents (or comparable agent-orchestration frameworks) for building self-evolving, autonomous QA agents, and drive the go/no-go decision on adoption
Collaborate with QA, Data Science, and Engineering to integrate AI-driven testing into the existing GitLab CI/CD pipeline (dev → test → staging → prod), alongside Terraform-provisioned, EKS/AgentCore-hosted services
Build resilient and adaptive automation that can detect and respond to changes in agent behavior, Bedrock Guardrails configuration, and routing logic
Develop data-driven quality analytics, including root-cause analysis, quality trends, and intelligent test prioritization, leveraging existing observability and evaluation data from OpenTelemetry, CloudWatch, X-Ray traces, and LangFuse offline evaluation runs
Lead research into emerging AI testing methodologies and mentor engineers through code reviews, workshops, and architectural guidance
BSc/MSc in Computer Science, AI, or a related discipline
6+ years of hands-on experience in QA/Test Automation, with strong experience designing and maintaining automation frameworks
1+ years of experience applying AI/ML in software testing or QA process improvement, with a proven track record of bringing AI/ML solutions into production workflows
Strong Python skills, with experience working in Python/AWS-centric technology environments; Java and/or TypeScript is a plus
Hands-on experience with LLM evaluation techniques, including LLM-as-a-judge, human-in-the-loop evaluation, RAG, and multi-agent orchestration patterns
Practical experience with DeepEval or comparable evaluation frameworks, along with familiarity with agent-orchestration SDKs such as Strands Agents
Knowledge of LLM tooling, vector databases, and MLOps pipelines
Experience integrating AI tooling into enterprise CI/CD environments, particularly GitLab, and working with containerized cloud-native environments such as Docker and Kubernetes/EKS
Strong communicator, able to influence technical decisions and drive engineering standards across teams
Experience with autonomous QA agents or agentic orchestration frameworks for self-evolving test suites
Experience with LLM observability tools such as LangFuse, LangSmith, or Arize, particularly for measuring probabilistic and adversarial robustness
Knowledge of AI ethics, fairness, and bias detection for guardrail and model validation
Experience with gRPC, WebSockets, and/or HTTP/2
Experience with AWS Bedrock, plus familiarity with GCP Vertex AI and/or Azure AI