UN
Technical Lead, AI Benchmarking and Evaluation Research (Remote, ROU)
- Hiring from
- Romania
- Work type
- Remote
- Posted
Is this job info correct?
511,036 remote jobs, straight from company career pages
100% free · New jobs every hour
Show job descriptionHide job description
Undelucram.ro on behalf of:
Crowdstrike SRL
About The Role
The CrowdStrike OCAIO team is looking for a senior technical lead to own how we measure AI models and agentic systems that perform cybersecurity tasks. This is a hands-on leadership role. You will lead a small team of researchers and you will personally help design, build and defend the benchmarks they ship.
Our mission is to establish rigorous, reproducible standards for how well AI and agentic systems perform real security work: malware analysis, reverse engineering, threat intelligence reasoning, alert triage, investigation and response. We build our ground truth with security subject matter experts, not with other models, and we score agents on what they actually did, not only on what they said.
As the Senior Lead, you define what "good" looks like for an AI system operating in a security workflow, you build the datasets, rubrics and judges that encode that definition, and you run the standing scorecards that engineering, product and research teams use to decide what ships. You set the technical bar and you meet it yourself.
What You'll Do
Hands-on benchmark work (about half of your time):
Crowdstrike SRL
About The Role
The CrowdStrike OCAIO team is looking for a senior technical lead to own how we measure AI models and agentic systems that perform cybersecurity tasks. This is a hands-on leadership role. You will lead a small team of researchers and you will personally help design, build and defend the benchmarks they ship.
Our mission is to establish rigorous, reproducible standards for how well AI and agentic systems perform real security work: malware analysis, reverse engineering, threat intelligence reasoning, alert triage, investigation and response. We build our ground truth with security subject matter experts, not with other models, and we score agents on what they actually did, not only on what they said.
As the Senior Lead, you define what "good" looks like for an AI system operating in a security workflow, you build the datasets, rubrics and judges that encode that definition, and you run the standing scorecards that engineering, product and research teams use to decide what ships. You set the technical bar and you meet it yourself.
What You'll Do
Hands-on benchmark work (about half of your time):
- Design and build benchmark datasets for cybersecurity agent tasks, from task definition through expert-labeled ground truth, scoring rubrics and release.
- Build and calibrate judges (LLM-based and programmatic) against expert labels, measure agreement and publish where they disagree.
- Build reproducible evaluation harnesses that capture agent actions and traces, not only transcripts, so results are defensible and comparable across model versions.
- Run and publish standing scorecards of CrowdStrike agents and frontier models on the same tasks, with protection, correctness and usefulness reported together.
- Investigate failure modes personally: where agents collapse, why, and what evidence proves it.
- Write the research: internal reports and external papers or talks that establish the benchmark's credibility.
- Lead, mentor and grow a team of researchers and engineers; review their designs, labels and code.
- Set the strategy, roadmap and success metrics for evaluating AI and agentic systems across security use cases.
- Define and enforce the evaluation methodology and the shared task and trace formats used across benchmarking and red teaming.
- Secure and coordinate subject matter expert labeling time from SOC, threat research and malware analysis teams, and set the quality bar for every label.
- Collaborate with engineering, product and threat research to turn evaluation findings into agreed pass/fail criteria and shipped improvements.
- Communicate results, trade-offs and limitations clearly to technical and executive audiences. Report what was not tested as plainly as what was.
- Hands-on cybersecurity expertise in at least one of: malware analysis and reverse engineering, incident response and threat hunting, detection engineering, SOC operations. You can label ground truth yourself and judge whether an expert's label is right.
- Demonstrated experience building evaluations, benchmarks or datasets for AI or LLM systems, with shareable artifacts (papers, repositories, internal benchmarks with measured adoption).
- Strong Python and data tooling skills. You write production-quality evaluation code, harnesses and analysis, and you review others' code.
- Practical experience with LLMs and agentic systems: tool use, multi-step planning, sandboxing, and how to instrument and trace agent behavior.
- At least 2 years leading technical people (as a team lead, tech lead or manager), with a record of growing researchers and engineers while remaining hands-on.
- Ability to define, design and standardize evaluation methodologies and reproducible testing pipelines, and to defend them under scrutiny.
- Broad knowledge of the cybersecurity landscape: attack techniques, defensive controls, and the analyst workflows that AI is meant to augment.
- Exceptional written and spoken communication. You can present measured findings and their limitations to executives and to researchers.
- Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes.
- Published research or talks on AI evaluation, LLM judging, or AI security.
- Experience with adversarial testing or red teaming of AI systems, and with scoring agents from observed side effects rather than text.
- Experience building LLM judges and measuring judge agreement with human experts.
- Experience with SOC platforms, SIEM/SOAR tooling and detection engineering.
- Understanding of MITRE ATT&CK and how it maps to detection and response workflows.
- Relevant security certifications (GREM, GCFA, GCIH, GCIA, OSCP or equivalent).
- Experience fine-tuning or distilling models on expert-labeled decision data.