Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
WA

LLM Red Team Specialist - Failure Modes & Edge Cases

Weekday AI
Posted 1 hour ago
🇺🇸United States🏠Remote💰$60.0–$90.0/hr📁Data & Analytics
Is this job info correct?

This role is for one of our clients Compensation: $60-$90 per hour Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red-teaming environment, you will design challenging, multi-step tasks that expose hidden vulnerabilities, reasoning gaps, and edge cases that traditional evaluations often miss. In this role, you'll collaborate closely with AI researchers to transform discovered failure modes into high-quality benchmark tasks that improve the robustness, safety, and reasoning capabilities of state-of-the-art AI systems. This is a fully remote, full-time engagement requiring approximately 35 hours per week . Key Responsibilities Investigate how frontier AI models perform across coding, machine learning, analytical reasoning, and complex problem-solving tasks. Identify hidden failure modes, edge cases, reasoning errors, and vulnerabilities that may not be apparent through standard testing. Design challenging evaluation tasks that accurately measure AI capabilities while remaining objective and reproducible. Document findings with clear technical explanations, supporting evidence, and reproducible methodologies. Collaborate with benchmark designers and AI researchers to refine evaluation tasks, eliminate loopholes, and strengthen grading criteria. Share insights and recommendations with cross-functional teams to continuously improve AI evaluation quality and benchmark coverage. Required Qualifications Master's degree, PhD, or equivalent practical experience in a STEM discipline involving research, coding, or advanced data analysis. Minimum 1 year of experience in AI research, research engineering, security research, AI evaluation, or a related technical field. Demonstrated experience identifying vulnerabilities, adversarial behaviors, edge cases, or failure modes in Large Language Models or other machine learning systems. Strong proficiency in Python and Git , with the ability to build custom scripts for experimentation, testing, and analysis. Solid understanding of modern Large Language Models, their strengths, limitations, and evaluation methodologies. Experience with AI benchmarking, model evaluation, adversarial testing, prompt engineering, or dataset creation is highly desirable. Excellent analytical thinking, creativity, and attention to detail, with the ability to solve ambiguous, open-ended problems independently. Outstanding written communication skills for documenting technical findings clearly and accurately. Ability to commit approximately 35 hours per week on a consistent basis. Preferred Qualifications Experience with AI safety, red teaming, adversarial machine learning, or security research. Background in benchmark design, evaluation framework development, or AI quality assurance. Experience creating reproducible technical experiments and documenting complex failure analyses. Familiarity with frontier AI research methodologies and model capability assessments. Why Join Help shape the future of AI evaluation by identifying critical weaknesses before they reach production. Work on cutting-edge AI systems alongside researchers developing next-generation language models. Apply your technical expertise to improve AI reliability, reasoning, and robustness. Contribute directly to benchmark development that influences the evolution of advanced AI technologies. Enjoy the flexibility of a fully remote engagement while working on impactful research initiatives. Equal Opportunity We are committed to fostering an inclusive and diverse environment where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process. Contract & Engagement Details Independent contractor engagement. Fully remote with flexible working hours. Expected commitment of approximately 35 hours per week . Project duration may be extended, shortened, or concluded based on project requirements and individual performance. Work does not require access to confidential or proprietary information from any current or former employer. Payments are issued weekly based on approved work completed. At this time, we are unable to support H1-B or STEM OPT candidates.

Similar jobs

Similar jobs

Sedgwick logo

Workers Compensation Claims Adjuster | NY Jurisdictional Knowledge & NY Licensing Required | Dedicated Public Entity Client & Capped Caseloads

Sedgwick

🇺🇸United StatesJun 16, 2026, 8:42 PM UTC
WA

Data Science & Quantitative Analysis Expert

Weekday AI

🇺🇸United States1 hour ago
WA

AI Rater Guidelines Writer (Linguist / Instructional Designer)

Weekday AI

🇺🇸United States1 hour ago
OpenDataJobs logo

AI Systems Engineer

OpenDataJobs

🇺🇸United States1 hour ago
Vsp logo

Senior Business Data Analyst

Vsp

🇺🇸United States3 hours ago
Crowdstrike logo

Sr. Data Scientist, Applied AI/ML

Crowdstrike

🇺🇸United States3 hours ago