CrowdStrike logo

Senior Technical Lead, AI Benchmarking and Evaluation Research (Remote, ROU)

Hiring from
Romania
Work type
Remote
Posted
Is this job info correct?
Show job description

As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn’t changed — we’re here to stop breaches, and we’ve redefined modern security with the world’s most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We're proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We’re always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you.



About the Role


The CrowdStrike OCAIO team is looking for a senior technical lead to own how we measure AI models and agentic systems that perform cybersecurity tasks. This is a hands-on leadership role. You will lead a small team of researchers and you will personally help design, build and defend the benchmarks they ship.


Our mission is to establish rigorous, reproducible standards for how well AI and agentic systems perform real security work: malware analysis, reverse engineering, threat intelligence reasoning, alert triage, investigation and response. We build our ground truth with security subject matter experts, not with other models, and we score agents on what they actually did, not only on what they said.


As the Senior Lead, you define what "good" looks like for an AI system operating in a security workflow, you build the datasets, rubrics and judges that encode that definition, and you run the standing scorecards that engineering, product and research teams use to decide what ships. You set the technical bar and you meet it yourself.


What You'll Do


Hands-on benchmark work (about half of your time):

- Design and build benchmark datasets for cybersecurity agent tasks, from task definition through expert-labeled ground truth, scoring rubrics and release.

- Build and calibrate judges (LLM-based and programmatic) against expert labels, measure agreement and publish where they disagree.

- Build reproducible evaluation harnesses that capture agent actions and traces, not only transcripts, so results are defensible and comparable across model versions.

- Run and publish standing scorecards of CrowdStrike agents and frontier models on the same tasks, with protection, correctness and usefulness reported together.

- Investigate failure modes personally: where agents collapse, why, and what evidence proves it.

- Write the research: internal reports and external papers or talks that establish the benchmark's credibility.


Leading the team and the function:

- Lead, mentor and grow a team of researchers and engineers; review their designs, labels and code.

- Set the strategy, roadmap and success metrics for evaluating AI and agentic systems across security use cases.

- Define and enforce the evaluation methodology and the shared task and trace formats used across benchmarking and red teaming.

- Secure and coordinate subject matter expert labeling time from SOC, threat research and malware analysis teams, and set the quality bar for every label.

- Collaborate with engineering, product and threat research to turn evaluation findings into agreed pass/fail criteria and shipped improvements.

- Communicate results, trade-offs and limitations clearly to technical and executive audiences. Report what was not tested as plainly as what was.


What You'll Need


- Hands-on cybersecurity expertise in at least one of: malware analysis and reverse engineering, incident response and threat hunting, detection engineering, SOC operations. You can label ground truth yourself and judge whether an expert's label is right.

- Demonstrated experience building evaluations, benchmarks or datasets for AI or LLM systems, with shareable artifacts (papers, repositories, internal benchmarks with measured adoption).

- Strong Python and data tooling skills. You write production-quality evaluation code, harnesses and analysis, and you review others' code.

- Practical experience with LLMs and agentic systems: tool use, multi-step planning, sandboxing, and how to instrument and trace agent behavior.

- At least 2 years leading technical people (as a team lead, tech lead or manager), with a record of growing researchers and engineers while remaining hands-on.

- Ability to define, design and standardize evaluation methodologies and reproducible testing pipelines, and to defend them under scrutiny.

- Broad knowledge of the cybersecurity landscape: attack techniques, defensive controls, and the analyst workflows that AI is meant to augment.

- Exceptional written and spoken communication. You can present measured findings and their limitations to executives and to researchers.

- Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes.


Bonus Points


- Published research or talks on AI evaluation, LLM judging, or AI security.

- Experience with adversarial testing or red teaming of AI systems, and with scoring agents from observed side effects rather than text.

- Experience building LLM judges and measuring judge agreement with human experts.

- Experience with SOC platforms, SIEM/SOAR tooling and detection engineering.

- Understanding of MITRE ATT&CK and how it maps to detection and response workflows.

- Relevant security certifications (GREM, GCFA, GCIH, GCIA, OSCP or equivalent).

- Experience fine-tuning or distilling models on expert-labeled decision data.




Benefits of Working at CrowdStrike:

  • Market leader in compensation and equity awards
  • Comprehensive physical and mental wellness programs
  • Competitive vacation and holidays for recharge
  • Paid parental and adoption leaves
  • Professional development opportunities for all employees regardless of level or role
  • Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections
  • Vibrant office culture with world class amenities
  • Great Place to Work Certified™ across the globe

CrowdStrike is proud to be an equal opportunity employer. We are committed to fostering a culture of belonging where everyone is valued for who they are and empowered to succeed. We support veterans and individuals with disabilities through our affirmative action program.


CrowdStrike is committed to providing equal employment opportunity for all employees and applicants for employment. The Company does not discriminate in employment opportunities or practices on the basis of race, color, creed, ethnicity, religion, sex (including pregnancy or pregnancy-related medical conditions), sexual orientation, gender identity, marital or family status, veteran status, age, national origin, ancestry, physical disability (including HIV and AIDS), mental disability, medical condition, genetic information, membership or activity in a local human rights commission, status with regard to public assistance, or any other characteristic protected by law. We base all employment decisions--including recruitment, selection, training, compensation, benefits, discipline, promotions, transfers, lay-offs, return from lay-off, terminations and social/recreational programs--on valid job requirements.


If you need assistance accessing or reviewing the information on this website or need help submitting an application for employment or requesting an accommodation, please contact us at recruiting@crowdstrike.com for further assistance.


About Us

CrowdStrike was founded in 2011 to fix a fundamental problem: The sophisticated attacks that were forcing the world’s leading businesses into the headlines could not be solved with existing malware-based defenses. Founder George Kurtz realized that a brand new approach was needed — one that combines the most advanced endpoint protection with expert intelligence to pinpoint the adversaries perpetrating the attacks, not just the malware.

There’s much more to the story of how Falcon has redefined endpoint protection but there’s only one thing to remember about CrowdStrike: We stop breaches.

Similar jobs

Apply for this job