About Bolter Bolter is an AI agent platform built to help people get real work done, without needing to stitch together multiple tools or spend weeks setting things up. You describe what you need in plain language and Bolter creates and runs agents that can carry out the work end to end. They retain context, remember how you work and can also create real, shareable apps such as trackers and dashboards. Bolter is funded and currently at the validation-sprint stage, working with a small group of high-impact operators. The product already exists. We are now looking for a founding designer to lead a significant redesign and establish the design foundations for what comes next. About the role As Bolter's AI Researcher, your job is to make agents that do real work trustworthy, capable, and useful - and to figure out how before we build it at scale. This is applied research with direct product impact. You design and run experiments, evaluate what works and what doesn't, and turn findings into decisions the engineering team ships. You'll work directly with the GM and the product engineer in a lean team, with specialised AI agents supporting implementation. The research problem goes beyond improving a model. It is working out how an agent that retains context, carries out multi-step work and creates software itself can be relied on to do it correctly. What you’ll do Design and run experiments to test how Bolter's agents behave - reliability, context retention, multi-step task completion, and where they fail. Build and own the evaluation layer. Design evals that measure whether agents actually do the work, not just whether they sound plausible. Research the frontier. Keep Bolter current on the state of the art in LLM agents, tool use, and reliability - and translate that into what we should build. Turn findings into decisions. You produce clear, actionable recommendations the engineering team can ship, not just papers. Prototype research into product. Take promising ideas from experiment to working prototype, and hand off what proves useful. Shape the research roadmap as Bolter grows. Why you're made for this A track record of applied AI/ML research - ideally in LLM application design, agentic systems, evals, or production AI reliability. Strong experimental design and analysis. You know how to test a hypothesis rigorously and read the result honestly. Strong engineering fundamentals - you can prototype your own experiments, not just direct others to run them. Deep familiarity with LLMs, tool use, context management, and the failure modes of agentic systems. Strong written communication. Research that isn't understood and acted on doesn't help. You're rigorous but pragmatic. You know when a finding is strong enough to act on and when it needs more evidence. You're comfortable with ambiguity. Many of the problems we're solving don't have textbook answers. You can leverage AI agents as a force multiplier - we run lean, with a small human core augmented by a fleet of specialised agents. You care about craft, but you ship. Bonus points Experience with LLM evals, hallucination mitigation, or production AI reliability at scale. Experience building or evaluating agentic systems, tool use, or autonomous workflows . A public track record - papers, open source, blog posts, side projects. We love researchers who ship outside work too. Experience in product-led research - where the output is a shipped feature, not just a finding. While we think the above experience could be important, we're keen to hear from people who believe they have valuable experience to bring to the role. If you identify with the team and mission, but not all of our requirements, then please still apply! Improbable Candidate Privacy Policy
Freelance Agent Evaluation Engineer
Toloka Ai
Founding Customer Success Manager - Bolter
Improbable
Product Engineer - Bolter
Improbable
Head of Design, AI-First
Puffy
University Admissions Tutor & Advisor
Brainlyne
UN Women - HeForShe Barbershop Facilitator (Consultant Roster)
UNDP