Remote | Humanities Evaluation Specialist — $40–$90/hour
24-MAGWe are sharing a specialised consulting opportunity for experienced humanities scholars, researchers, and professionals with strong expertise in interpretation, argumentation, source analysis, and analytical writing to contribute to an advanced AI evaluation project.
Selected professionals will develop original, high-difficulty humanities questions, research and verify authoritative answers, and test whether advanced AI systems can reason effectively across interpretation, context, argument, and complex source material. The work requires strong analytical judgement, source triangulation, written precision, and the ability to create rigorous evaluation material that goes well beyond surface-level recall. No prior experience in AI is required.
Key Responsibilities
Humanities Question Development
- Develop original, high-difficulty question-and-answer pairs across humanities disciplines
- Create prompts that test interpretation, argumentation, contextual understanding, and analytical reasoning
- Design questions that require deep insight rather than straightforward factual recall
- Ensure tasks reflect genuine disciplinary complexity and expert-level judgement
- Maintain originality, intellectual depth, and methodological rigour across submitted work
Research & Source Triangulation
- Research and verify answers using primary texts, authoritative references, and reputable secondary sources
- Provide clear citations supporting factual claims, interpretations, and analytical conclusions
- Triangulate information across multiple sources where necessary
- Distinguish well-supported interpretations from weak, ambiguous, or insufficiently evidenced claims
- Document reasoning and source support clearly and consistently
AI Model Testing & Evaluation
- Test developed questions against advanced AI models
- Identify questions that are insufficiently challenging or susceptible to superficial answers
- Iterate on prompts to increase difficulty while preserving accuracy and clarity
- Analyse where AI systems succeed or fail in humanities reasoning and interpretation
- Refine question structure to better measure genuine understanding and analytical capability
Written Precision & Quality Review
- Write questions and answers with exceptional clarity and attention to detail
- Eliminate ambiguity while preserving the intended intellectual difficulty
- Produce well-supported answers that clearly explain the reasoning behind conclusions
- Incorporate reviewer feedback and revise submissions where required
- Follow project guidelines and maintain consistent quality standards across deliverables
Ideal Profile
- Advanced academic background or substantial professional experience in a humanities discipline
- Relevant fields may include history, literature, philosophy, linguistics, classics, art history, journalism, or related areas
- Strong analytical writing and close-reading skills
- Demonstrated ability to formulate nuanced, challenging, and original questions
- Experience researching and triangulating information from authoritative sources
- Strong ability to document reasoning and citations clearly
- Exceptional attention to detail and written precision
- Self-directed and reliable when completing expert-level work independently
- Comfortable working within remote, asynchronous project environments
- Familiarity with AI model evaluation or training is advantageous but not required
Engagement Details
- Independent contractor engagement
- Fully remote
- Compensation: $40–$90/hour
- Compensation is output-based, with payment made for tasks that meet project specifications
- Minimum weekly submission requirements apply
- Work will involve humanities question development, source research, reference-answer authoring, AI model testing, and iterative refinement
- Task completion time may vary depending on research complexity, discipline, and individual workflow
- Selected professionals should be prepared to begin their first tasks within approximately 24–48 hours of completing onboarding
- Roles are typically filled within approximately 48 hours
- Project scope, workload, and evaluation standards may evolve depending on project requirements
- Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, research group, or other third party
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy