Senior QA Engineer
- Hiring from
- United States
- Work type
- Remote
- Posted
Show job descriptionHide job description
Description
Vi Operate puts AI agents into the daily operations of large health systems, providers, and pharma organizations. Augment is the voice capability underneath it: the agents that call patients and providers, hold a real conversation, and take action in the systems where the work actually lives. This is a dedicated quality role for that work.
Testing a voice agent is not like testing software that returns the same answer twice. A similar call can go three different ways. A model update can quietly change behavior that nothing in a unit test would catch. When a call goes badly, the cause could be the speech recognition, the model, the voice synthesis, the telephony layer, or the integration underneath, and the transcript alone won't tell you which. Making that testable, repeatable, and catchable before a client hears it is the job.
You will own quality for agents that call real patients, providers, and caregivers. That means building the test and evaluation systems our engineers develop against, deciding what has to pass before something ships, and being the person who can take a bad call recording and say what actually went wrong. We need someone who has done classical QA properly and has since tested voice and LLM systems in production, because this role needs all three.
What You'll Own
- Voice testing and debugging. Call recordings, transcripts, latency traces, and logs. Isolating which layer failed: speech to text, the model, text to speech, telephony, or the integration behind it. Turning a vague "the call went badly" into a reproducible ticket an engineer can act on.
- Evaluation and regression testing. Golden test sets, simulated calls, judged transcripts, accuracy on classification and extraction, and structured output validation. Regression coverage that still holds when the underlying model changes.
- Testing things that aren't deterministic. Deciding what "passing" means when the output varies run to run. Setting tolerances, handling flake honestly rather than by rerunning until green, and keeping the suite trustworthy enough that engineers act on a failure.
- Release criteria. What has to be true before an agent goes live for a client, and what has to stay true afterward. You set the bar and you hold it, including when a deployment date is pushing on it.
- Classical QA for the rest of the product. Test plans, regression suites, bug triage, reproductions, and CI integration for everything around the agent.
- Quality in production. Sampling live calls after release, spotting drift and degradation, and getting it in front of the right engineer before the client raises it.
- Breaking it on purpose. Red-teaming the agent before someone outside does. Prompt injection through the documents it reads, callers who talk it off script or push it toward giving medical advice, people claiming to be someone they aren't to get patient information out of it. You test that the guardrails and access controls actually hold under pressure, not just that they exist.
What We're Looking For
- QA on production voice agents. You have tested a system that talks to people on the phone, and you can walk through a real failure you diagnosed: what you heard, where you looked, and which layer turned out to be at fault. This is required and it's what we screen hardest on.
- QA on LLM systems. You have tested something built on a language model in production, not as a side project, and you can describe how you decided whether it was working.
- Real classical QA depth. Test strategy, test planning, regression suites, bug triage, and release process. You have worked somewhere QA had teeth.
- Testing non-deterministic systems. You have dealt with output that varies and made it testable anyway, and you can explain the tradeoffs you made to get there.
- Python and test automation. You write your own tooling and wire it into CI rather than filing a ticket for someone else to.
- Judgment about severity. You can tell the difference between a cosmetic defect and something that will mislead a patient on a phone call, and you can defend that call to an engineer who disagrees.
- A system mindset. You look at the whole system rather than one piece, and your first instinct is to find where it comes apart. You think about the seams between components, the inputs nobody planned for, and the person on the other end trying to get something out of it that they shouldn't.
Nice To Have
- Background in IVR, contact center, or CCaaS quality (Genesys, NICE, Five9, Amazon Connect, Twilio, or similar)
- Audio and network condition testing: codecs, jitter, packet loss, barge-in and interruption handling
- Familiarity with evaluation tooling such as LLM-as-judge harnesses, prompt regression frameworks, or simulated conversation testing
- Healthcare or pharma experience, and comfort working under HIPAA
- Experience standing up a QA function rather than joining one
- Load and performance testing on real-time systems
- Spanish proficiency, enough to listen to a call and judge whether the agent handled it