LUZID QA Engineer São Paulo, in-person 4x/week | Full-time About Luzid Luzid is building the operating system for SAP delivery — the infrastructure layer underneath the most complex and expensive IT projects companies run. SAP implementations move billions of dollars and thousands of consultants, but the knowledge generated in these projects is never captured: it lives in people's heads, disappears when the project ends, and gets rebuilt from scratch on every new engagement. Luzid solves this with an agentic orchestration platform that coordinates every phase of SAP transformation, capturing decisions, tests, and context as a byproduct of work already underway — building an organizational memory that accumulates across projects instead of resetting. We're a team based in San Francisco and São Paulo, moving fast and already deployed with some of the largest SAP integrators in the world, including delaware Consulting, KPMG, and NTT Data. We have reached >$3M ARR in 11 months and raised $21M in funding. The team comes from Stanford, Meta, Two Sigma, Microsoft, and NASA. The Opportunity Luzid generates functional specifications, configuration, and test artifacts for SAP implementations. Our customers are consultants running multi-million dollar programs, and they act on what the platform produces. A confidently wrong output is worse than no output at all: it survives review, reaches a client deliverable, and is discovered at go-live. Preventing that is the job. So this role is two things at once. It is conventional software QA — automated end-to-end coverage, regression suites, CI discipline, release confidence. And it is evaluation engineering for non-deterministic systems: building the harnesses, golden datasets, and scoring criteria that tell us whether an agent's output is actually correct, and whether last week's prompt change quietly made it worse.. What You'll Do Own the automated test suite. Build and maintain end-to-end, integration, and regression coverage across the platform, and keep it fast and trustworthy enough that people actually act on a red build. Build the evaluation harness for agent output. Golden datasets, scoring criteria, and regression testing for LLM-generated artifacts — so we can tell a prompt or model change that improved things from one that broke them. Define the quality gates. Decide what blocks a release, write the criteria down, and hold the line when there is pressure to ship. Own CI reliability. Flaky tests, slow pipelines, and unclear failures are quality problems in their own right, and they are yours. Run exploratory and release testing on the paths automation does not reach, particularly the customer-facing workflows consultants use under deadline pressure. Triage and drive defects to closure with reproducible cases, clear severity calls, and follow-through until the fix is verified. Feed quality signal back into the product. You will see the failure patterns before anyone else; that belongs in the roadmap, not just in a bug tracker. What We're Looking For Required 3+ years in QA, SDET, or test engineering with real ownership of automated coverage for a production product. Strong test automation skills in Python and/or TypeScript, including a modern E2E framework (Playwright, Cypress, or equivalent) and API-level testing. CI/CD fluency. You have built and maintained pipelines, and you have opinions about what belongs in one. Comfort testing systems without deterministic outputs. Either direct experience evaluating LLM or ML output, or a clear, credible approach to how you would do it. Judgment about severity. You can tell the difference between a cosmetic defect and one that will cost a customer their go-live date, and you prioritize accordingly. AI native. You use AI tooling in your daily workflow and treat it as a force multiplier rather than a shortcut. High agency in an environment with no playbook. You will be defining the function, not staffing it. Professional-to-fluent English. Prior work at an early-stage company where QA was something you created rather than joined. Nice to have Experience evaluating LLM-powered or agentic features in production: eval design, hallucination detection, human-in-the-loop review workflows. Familiarity with enterprise software ecosystems — SAP, Salesforce, ERP or CRM more broadly — or willingness to learn the domain deeply. SAP testing tooling exposure (Tricentis, Worksoft, or comparable), or experience with SAP test cycles and UAT. Performance, load, or security testing experience. You write code well enough to fix the bug, not just file it. Why Join Luzid Now Define the function. You are building the quality bar for the company, not maintaining someone else's suite. A genuinely hard testing problem. Non-deterministic output, enterprise-grade consequences, and no established playbook for either. Real stakes. Your work sits between an agent's output and a multi-million dollar client deliverable. Direct access to engineering, product, and the founding team. Competitive compensation with meaningful equity.
Mobile Staff QA Engineer I
Housecall Pro
QA Engineer II
Housecall Pro
Senior Automation QA Engineer
Ciklum
Senior QA Engineer (FX Functional Test)
Exadel Inc (Website)
Senior QA Engineer
Exadel
Data QA Engineer – Enterprise Data
Truelogic