24-MAG logo

Remote | Research Evaluation Specialist (PhD / Researcher / Professor) — $40–$90/hour

24-MAG
Posted 7 hours ago
United StatesRemote$40–$90/hrResearch & Science
Is this job info correct?

We are sharing a specialised consulting opportunity for PhD-level researchers, professors, and advanced subject-matter experts to contribute to an AI evaluation project focused on developing rigorous, research-grade test material for advanced AI systems.

Selected professionals will create original, high-difficulty questions within their areas of expertise, develop thoroughly sourced reference answers, and test whether advanced AI systems can solve problems requiring genuine specialist reasoning. The work requires strong research judgement, source triangulation, methodological precision, and the ability to produce defensible, unambiguous evaluation material. No prior experience in AI is required.

Key Responsibilities

Advanced Research Question Development

  • Create original, high-difficulty question-and-answer pairs at the frontier of your academic or research discipline
  • Design questions requiring advanced reasoning, methodological nuance, or synthesis across multiple sources
  • Develop problems that cannot be solved reliably through superficial pattern matching or simple lookup
  • Ensure questions reflect genuine specialist-level knowledge and research complexity
  • Maintain originality, technical depth, and intellectual rigour across submitted tasks

Source Verification & Research Triangulation

  • Research and verify answers using primary sources and authoritative references
  • Triangulate information across multiple high-quality sources when necessary
  • Provide clear citations and supporting reasoning for reference answers
  • Distinguish well-supported conclusions from uncertain, contested, or insufficiently evidenced claims
  • Ensure all supporting material is accurate, traceable, and appropriate for expert-level evaluation

AI Model Testing & Difficulty Calibration

  • Test developed questions against advanced AI systems
  • Identify questions that are insufficiently challenging or vulnerable to shortcut solutions
  • Iteratively increase difficulty while preserving factual accuracy and methodological validity
  • Analyse where AI systems succeed or fail when addressing specialist research problems
  • Refine questions to better measure genuine reasoning and subject-matter capability

Precision, Review & Quality Standards

  • Write questions and answers with exceptional clarity and precision
  • Eliminate ambiguity that could prevent objective or defensible evaluation
  • Incorporate reviewer feedback and revise deliverables accordingly
  • Follow project guidelines, formatting requirements, and quality standards consistently
  • Maintain reliable quality across independent, remote research workflows

Ideal Profile

  • Completed PhD, active PhD candidacy, or equivalent advanced research experience as a specialist, researcher, or professor
  • Demonstrated record of scholarly research and deep subject-matter expertise
  • Strong familiarity with primary literature and authoritative research sources within your field
  • Excellent analytical reasoning and ability to engage with methodologically complex problems
  • Experience sourcing, verifying, and triangulating information across authoritative references
  • Exceptional attention to detail and written precision
  • Strong written English communication skills
  • Ability to create challenging, original, and methodologically sound research questions
  • Proven self-direction and reliability when completing expert-level work independently
  • Comfortable incorporating reviewer feedback and iterating on technical material
  • Previous experience in AI training, model evaluation, or related research workflows is advantageous but not required

Engagement Details

  • Independent contractor engagement
  • Fully remote
  • Compensation: $40–$90/hour
  • Compensation is output-based, with payment made for tasks that meet project specifications
  • Minimum weekly submission requirements apply
  • Work will involve advanced research question development, source triangulation, reference-answer authoring, AI model testing, and iterative difficulty calibration
  • Task completion time may vary depending on discipline, research complexity, and individual workflow
  • Selected professionals should be prepared to begin their first tasks within approximately 24–48 hours of completing onboarding
  • Roles are typically filled within approximately 48 hours
  • Project scope, workload, and evaluation standards may evolve depending on project requirements
  • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, research group, or other third party

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy

Similar jobs