About the Role We're looking for a Clinical Research Scientist to help build the next generation of evaluations for AI systems used in mental health and other clinically sensitive settings. As large language models become part of how people seek advice, emotional support, and health information, there is a growing need for rigorous ways to understand how these systems behave in real-world conversations. Many of the most important questions including how models respond to psychological distress, uncertainty, or vulnerable users, can't be answered with traditional AI benchmarks alone. They require clinical expertise, careful study design, and realistic evaluations grounded in human behavior. You'll work with researchers and engineers to design clinician-informed benchmarks, develop new evaluation methodologies, and build datasets that measure model behavior in realistic, multi-turn interactions. The role combines clinical research, behavioral science, and AI evaluation, with opportunities to publish, collaborate with leading universities and help shape emerging standards for evaluating AI. We welcome applicants from academia, hospitals, nonprofit research institutes, and digital health organizations who are excited to bring their research into industry while continuing to publish and collaborate with the broader research community. What You'll Do Design clinician-informed benchmarks, datasets, and evaluation methods for AI systems used in mental health and other clinically sensitive domains. Design and conduct validation studies to ensure benchmark performance reflects real-world model behavior. Partner with psychologists, psychiatrists, researchers, and academic collaborators to identify important evaluation problems and translate them into rigorous benchmarks. Build collaborative research projects with universities, hospitals, and nonprofit organizations. Analyze model behavior, publish research findings, and communicate results through technical reports and presentations. Collaborate with research engineers to implement large-scale evaluation pipelines and benchmark infrastructure. Help shape our research agenda in mental health AI evaluation and identify emerging research directions. Requirements Research background: PhD, PsyD, MD, or equivalent research experience in Clinical Psychology, Psychiatry, Behavioral Science, Public Health, or a related field. Research experience: Demonstrated experience designing and leading research projects, including study design, data collection, statistical analysis, and scientific writing. Publications: Track record of publishing independent research in peer-reviewed journals or conferences. Collaborative research: Demonstrated ability to initiate and lead collaborative research with external partners, including universities, hospitals, or other research organizations. Research methods: Strong understanding of behavioral research methods, human subjects research, survey design, psychometrics, qualitative or quantitative methods, or clinical study design. Communication: Excellent written and verbal communication skills, including experience presenting research to diverse audiences. Nice to Haves Research focused on adolescent mental health, suicide prevention, psychotherapy, digital mental health, clinical decision making, or related areas Experience studying how people interact with AI or other digital technologies Familiarity with large language models or AI evaluation Experience with longitudinal studies, conversation analysis, or real-world behavioral datasets Experience leading IRBs, multi-site studies, or collaborations across institutions Existing collaborations within academia or healthcare that you'd like to continue growing What We Offer Highly competitive salary and meaningful ownership. Excellence is well rewarded. Relocation and transportation support Health/dental insurance coverage Lunch and dinner provided, free snacks/coffee/drinks Unlimited PTO Opportunity to publish and present your work About Us Founding team : The core methodology behind this platform comes from NLP evaluation research we had done at Stanford. We raised a $5M seed from some of the top institutional and angel investors in the valley. Our team has prior work experience at NVIDIA, Meta, Microsoft, Palantir and HRT. Collectively, we have over 300 citations in our published work. Our early team include Stanford PhDs, ex-Jane Street quants, and the first designer at Snorkel. Tech stack : We use Python for most things at Vals. Our platform is built on Django, with a React frontend. All of the infra is on AWS using CDK for IaC. What We're Looking For Learning velocity: The role encompasses a wide variety of tasks. Rather than expecting you to be an expert on Day 1, we are looking for someone who can learn new skills and technologies extremely quickly. Ownership : Working in a small, talent-dense team, we expect everyone to show initiative to build where it's needed, not where it's asked. We strive for autonomy over consensus. This is especially true for this role. Intensity : The LLM landscape is constantly changing. Foundation model labs are continuously pushing the frontier. The unicorn companies that will emerge from this technology shift are being built now. Those that win will have an incredibly high speed of execution. Solution-oriented mindset : We're looking for people who see opportunities to craft solutions at each juncture, not those who pass hard problems to others or admit defeat. Further Reading: Hugging Face blog on evaluation Anthropic’s blog on challenges in evaluation New York Times article on issues in benchmarking Stanford HAI report showing hallucinations in legal tech tools
Boeing Summer 2027 Internship Program (Paid) – Environment, Health, and Safety (EHS)
Boeing
Boeing Summer 2027 Internship Program (Paid) – Environment, Health, and Safety (EHS)
Boeing
Chief Financial Officer (CFO) in Training - Southcoast Behavioral Health
Acadiahealthcare
AI Success Engineer - Healthcare & Life Sciences
OpenAI
Duke Health – General Neurologist at Main Campus
Dukehealth
Founding AI Engineer (Healthcare)
Clara