SF

Product Lead, AI Agent Evaluation & Research Products

Hiring from
United States
Work type
Remote
Posted
Sep 24, 2026
Is this job info correct?
We’re looking for a deeply technical product person to own the release of frontier AI products at the intersection of agents, evals, benchmarks, open-source tooling, and developer/community adoption.
This person will turn research ideas into products people can actually use: public benchmarks, agent competitions, evaluation platforms, datasets, developer tools, papers, launch campaigns, demos, and partner-facing programs. You’ll work closely with research, engineering, design, community, BD, and marketing to take ambiguous AI research directions from “interesting idea” to shipped product with users, metrics, documentation, launch narrative, and iteration loops.
You should be comfortable operating between product management, research translation, developer relations, and technical program execution. You should understand AI agents, evals, benchmarks, LLM tooling, and why measurement is becoming the core product surface for AI.
What You’ll Do
  • Own 0-to-1 product launches for AI agent and eval products, including benchmarks, leaderboards, datasets, developer challenges, research demos, and open-source releases.
  • Translate research agendas into concrete product specs, user journeys, launch plans, success metrics, and roadmap priorities.
  • Work with researchers to package technical work into usable artifacts: docs, examples, benchmarks, rubrics, eval harnesses, blog posts, landing pages, and demos.
  • Drive developer-facing programs like competitions, cohorts, hackathons, public challenges, and partner-backed benchmark launches.
  • Define the product experience around agent improvement loops: task definition, requirement mining, skill discovery, adaptation, evaluation, and performance improvement.
  • Coordinate engineering, design, research, community, marketing, and BD so releases ship with polish and momentum.
  • Identify target users, including AI developers, researchers, enterprise teams, open-source builders, and technical partners.
  • Collect feedback from users and partners, then turn it into product improvements, research questions, and follow-on releases.
  • Build launch narratives that make technical work legible and exciting without oversimplifying the research.
What We’re Looking For
  • 4+ years in product, technical product, founder, developer platform, AI tooling, or research-product roles.
  • Strong understanding of modern AI agents, evals, benchmarks, LLM workflows, and open-source/developer ecosystems.
  • Proven ability to ship ambiguous technical products from concept to launch.
  • Excellent writing and communication skills; can turn complex research into clear product positioning.
  • High agency and comfort working across research, engineering, GTM, community, and external partners.
  • Strong product taste: knows when a research artifact is not yet a product, and what it needs to become one.
  • Comfortable reading papers, discussing eval design, and working with technical users.
  • Prior startup experience is a plus
  • Bonus: experience with AI/ML infrastructure, LLM evals, Python, open-source launches, developer communities, hackathons, or benchmark design.
Success Looks Like
  • AI research releases ships on time with clear positioning, docs, demos, and user flows.
  • Developer challenges and benchmark launches attract strong builders and generate useful agent traces, submissions, feedback, and community energy.
  • Research artifacts become repeatable product surfaces rather than one-off announcements.
  • The team has a crisp roadmap connecting agent evals, automated improvement, developer adoption, and partner use cases.


Similar jobs

Apply for this job