We are sharing a specialised part-time consulting opportunity for psychologists, behavioural scientists, and academic researchers with advanced expertise in psychological theory, research methodology, psychometrics, human factors, or applied behavioural science. This role supports an AI research initiative focused on developing rigorous academic benchmarks in psychology. Selected experts will author or verify challenging multiple-choice assessment content, evaluate conceptual and empirical accuracy, develop well-supported solution rationales, and help establish high-quality reference standards for advanced AI systems. Key Responsibilities Psychology Question Authoring Create original multiple-choice questions that test deep conceptual understanding rather than surface-level recall Develop challenging problems within relevant areas of psychological expertise Ensure questions are self-contained, precise, unambiguous, and independently solvable Design questions appropriate for undergraduate, advanced undergraduate, and postgraduate difficulty levels Produce one correct answer alongside nine plausible and carefully constructed distractors Question Verification & Academic Review Review pre-written psychology questions for accuracy, clarity, rigour, and completeness Identify conceptual errors, ambiguity, unsupported assumptions, or solvability issues Edit questions and answer choices where required Verify that designated correct answers are supported by psychological theory and empirical evidence Document substantive changes and explain the reasoning behind them Psychological Research & Methodology Apply graduate-level psychological theory and empirical research to assessment design Evaluate study designs, experimental evidence, measurement approaches, and research conclusions Develop questions requiring interpretation of behavioural and psychological data Distinguish between competing theoretical explanations and methodological approaches Ensure assessment content reflects established research standards and current academic literature Psychometrics & Measurement Develop rigorous problems involving psychological measurement and assessment Apply concepts involving reliability, validity, scale construction, and measurement error Evaluate statistical and methodological assumptions underlying psychological instruments Create questions requiring interpretation of psychometric evidence Distinguish technically sound measurement approaches from plausible but flawed alternatives Human Factors & Applied Psychology Develop content involving human factors and engineering psychology Create problems addressing cognition, decision-making, attention, perception, and human-system interaction Apply psychological principles to technology, design, safety, and performance contexts Develop assessment content involving consumer and market psychology Evaluate behavioural reasoning in practical and applied settings Digital Health & Emerging Psychology Domains Develop or review content involving digital health and technology-mediated behavioural interventions Apply psychological research principles to emerging health and behavioural technologies Create assessment content relating to neuromorphic engineering where relevant to cognition and human-system interaction Develop academically rigorous questions involving psychedelic-assisted therapy Ensure emerging-domain content is grounded in reputable empirical research and appropriately framed Academic Research & Benchmark Quality Provide 1–5 academic references per question where required Use reputable sources such as peer-reviewed journals, academic publications, and university repositories Develop clear, structured solution explanations supporting the correct answer Rate question difficulty according to defined Medium, Hard, and Expert standards Help maintain consistent academic quality across benchmark content Contribute to gold-standard evaluation materials used to assess advanced AI capabilities Ideal Profile PhD, PsyD, or current doctoral candidacy in Psychology or a closely related discipline Master's degree may be considered for candidates with exceptional expertise in a relevant subdomain Strong command of graduate-level psychological theory, research methodology, and empirical literature Expertise in one or more areas including Human Factors, Engineering Psychology, Consumer Psychology, Market Psychology, Psychometrics, Digital Health, Psychedelic-Assisted Therapy, or related behavioural-science fields Strong understanding of experimental design, evidence evaluation, and psychological measurement Excellent written English and ability to communicate complex concepts clearly and precisely Academic research publications are highly valued Clinical licensure is advantageous for relevant areas of expertise University-level teaching, assessment design, or examination-development experience is highly valued Strong attention to conceptual precision, methodological validity, and evidence quality Engagement Details Part-time independent contractor engagement Fully remote and asynchronous Expected availability of at least 10 hours per week Flexible scheduling based on project requirements Compensation: $40–$55/hour Contributors may be assigned either question-authoring or question-verification workstreams Projects may be extended, shortened, or concluded based on project needs and performance Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party H1-B and STEM OPT support is unavailable for this engagement About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy .
Hourly Training Specialist – NAM Major Corrective-2
Gevernova
Hourly Training Specialist – NAM Major Corrective-2
Gevernova
Remote | Philosophy Research & AI Benchmark Specialist — $40–$55/hour
24-MAG
Remote | History & Political Science AI Benchmark Specialist — $35–$50/hour
24-MAG
Senior Patient Financial Services Specialist 24 hour TEMPORARY
Hebrewseniorlife
Inbound Sales Specialist - (Pacific Time Hours)
Trupanion