Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
2M

Remote | LLM Red Team & Benchmark Evaluation Specialist — $55–$85/hour

24 Mag
Posted 2 hours ago
🇺🇸United States🏠Remote💰$55.0–$85.0/hr📁Data & Analytics
Is this job info correct?

We are sharing a specialised full-time consulting opportunity for AI evaluation and research professionals with experience identifying failure modes, vulnerabilities, edge cases, and hidden weaknesses in large language models and machine learning systems. This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected professionals will probe complex model behaviour, design challenging multi-step tasks, document reproducible failures, and collaborate with researchers to strengthen benchmark quality across coding, machine learning, experimentation, and technical analysis. Key Responsibilities Adversarial Model Evaluation Probe frontier AI models across coding, machine learning, analysis, and multi-step agentic tasks Identify subtle errors, vulnerabilities, edge cases, and misleadingly plausible outputs Investigate situations where models appear capable while reaching incorrect or unsupported conclusions Design reproducible experiments to isolate and validate model failure modes Benchmark & Challenge Design Convert observed model weaknesses into rigorous benchmark tasks Develop challenges that are technically demanding while remaining fair and objectively assessable Define clear task requirements, expected outcomes, and evaluation criteria Ensure tasks require genuine reasoning rather than allowing shortcuts or superficial pattern matching Failure Analysis & Documentation Document findings with clear evidence, methodology, and reproducible steps Explain why a model failed and which capabilities or assumptions contributed to the error Produce detailed technical write-ups for researchers and task authors Track recurring failure patterns across models, prompts, and evaluation environments Task Strengthening & Research Collaboration Work with task authors to close loopholes, grading gaps, and unintended shortcuts Review benchmark tasks for ambiguity, exploitability, and evaluation reliability Share insights with researchers and other specialists to improve benchmark coverage Participate in iterative calibration, peer review, and task-refinement workflows Ideal Profile Strong candidates may have: At least 1 year of experience in research, research engineering, security, AI evaluation, or a related technical role Demonstrated experience identifying vulnerabilities, edge cases, or failure modes in LLMs or ML systems Background in red teaming, adversarial testing, security research, benchmark development, or rigorous model evaluation Working proficiency in Python and Git Ability to develop scripts, probes, and analyses independently Strong familiarity with LLM capabilities, limitations, and evaluation techniques Excellent written communication and technical documentation skills Creativity, precision, and persistence when working through ambiguous research problems Reliable availability for approximately 35 hours per week Educational Background A master's degree or PhD in a STEM field is highly relevant Equivalent practical experience in a research-intensive domain involving coding and data analysis may also be considered Academic or professional work involving machine learning, computer science, statistics, security, mathematics, or engineering may strengthen an application Publications, benchmark contributions, technical research, or impactful open-source work may also be valuable Nice to Have Experience in AI training, model evaluation, or benchmark authoring Background developing adversarial prompts or red-team evaluation suites Familiarity with agentic systems and multi-step tool-use evaluations Experience assessing coding, ML, or technical-analysis tasks Knowledge of experimental design and reproducibility Experience developing grading rubrics or automated evaluation methods Familiarity with security research or vulnerability assessment Prior collaboration with AI research or engineering teams Why This Opportunity Investigate where frontier AI models fail across complex technical tasks Help build stronger and more reliable agentic evaluation benchmarks Work directly with researchers on high-impact AI evaluation challenges Apply coding, experimentation, and analytical expertise to open-ended problems Contribute to stronger evaluation standards for advanced AI systems Participate in a structured full-time remote role with competitive hourly compensation Contract Details Full-time W-2 contingent employment opportunity Fully remote within the United States Expected commitment of approximately 35 hours per week Competitive rates between $55–$85 per hour depending on expertise and project scope Individual tasks may require one to two days of focused technical work Work may include model probing, benchmark design, failure analysis, technical documentation, and task refinement Close collaboration with research and benchmark-development teams Engagement scope and duration may evolve according to project requirements and performance About the Platform This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams. By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy .

Similar jobs

Similar jobs

2M

Remote | AI Safety Specialist (English & Norwegian) — $45–$60/hour

24 Mag

🇺🇸United States2 hours ago
2M

Remote | Licensed Civil Engineer & AI Evaluation Specialist — $60–$85/hour

24 Mag

🇺🇸United States2 hours ago
Mercor logo

AI Safety Specialist - Fully Remote | Upto $62/hr

Mercor

🌍Belgium, United Kingdom, United States2 hours ago
Mercor logo

AI Safety Specialist - Fully Remote | Upto $22/hr

Mercor

🌍Australia, Canada, India, Pakistan, Singapore, United Kingdom, United States2 hours ago
SV

Business Intelligence Senior Analyst, Information Technology

Svclnk

🇺🇸United States2 hours ago
Affiliates Commonspirit logo

RN Supervisor UM Prior Auth

Affiliates Commonspirit

🇺🇸United States1 hour ago