Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
WA

Technology AI Evaluation Expert

Weekday AI
Posted 2 weeks ago
🇺🇸United States🏠Remote💰$60.0–$75.0/hr📁Data & Analytics
Is this job info correct?

This role is for one of our clients Compensation: $60-$75 per hour Join an advanced AI research initiative focused on improving how next-generation AI systems understand professional documents, execute complex instructions, and reason through real-world technical workflows. We are seeking experienced technology professionals to design high-quality benchmark tasks that evaluate AI performance across software engineering and data science domains. In this role, you will create realistic, multi-step evaluation tasks based on technical documentation, code repositories, API references, architecture diagrams, and other workplace resources. Your work will help measure and improve the ability of AI models to interpret technical information, follow detailed instructions, and generate accurate, well-structured outputs. This is a fully remote, independent contractor opportunity with flexible working hours. Key Responsibilities Design AI Evaluation Tasks Create realistic, multi-step benchmark tasks based on professional technology workflows. Develop challenges using technical specifications, architecture documents, API documentation, codebases, web research, and code execution. Ensure each task includes a clearly defined expected output and objective evaluation criteria. Develop Evaluation Standards Write comprehensive ground-truth solutions and structured scoring rubrics. Design tasks that assess reasoning, technical understanding, instruction following, and output quality. Maintain high standards of technical accuracy, clarity, and reproducibility. Contribute Domain Expertise Apply real-world knowledge from software engineering, data science, or analytics to create authentic evaluation scenarios. Collaborate with research teams to improve benchmark quality and consistency. Continuously refine tasks based on project feedback and evolving evaluation requirements. Required Qualifications Minimum 3 years of hands-on professional experience in one or more of the following areas: Software Engineering Data Science Data Analytics Strong understanding of technical documentation, software development workflows, and engineering best practices. Experience working with codebases, APIs, technical specifications, or system architecture documentation. Excellent analytical thinking and problem-solving skills. Strong written communication with the ability to create clear technical instructions and evaluation criteria. Ability to work independently while maintaining high standards of accuracy and consistency. Engagement Details Independent contractor engagement. Fully remote with flexible working hours. Expected commitment of 15–20 hours per week . Projects may be extended, shortened, or concluded based on business needs and performance. Weekly payments processed through supported payment platforms. Why Join Help shape the next generation of AI systems for technical reasoning and document understanding. Work on intellectually challenging projects involving real-world engineering and data science workflows. Apply your technical expertise to improve advanced AI evaluation benchmarks. Enjoy flexible remote work with meaningful impact on AI research. Equal Opportunity Statement We are committed to providing equal opportunities to all qualified applicants without regard to legally protected characteristics. Reasonable accommodations are available upon request. Contract Information Independent contractor engagement. Fully remote work completed on your own schedule. Weekly payments are processed based on approved work completed. Work does not involve access to confidential or proprietary information from any employer, client, or institution. Please note that visa sponsorship is not available for this opportunity.

Similar jobs

Similar jobs

Volga Partners logo

Senior Investment Banking Subject Matter Expert (AI Evaluation) | U.S.

Volga Partners

🌍India, Pakistan, United States2 days ago
Volga Partners logo

Senior Quantitative Finance Subject Matter Expert (AI Evaluation) | U.S.

Volga Partners

🇺🇸United States2 days ago
Dexis logo

Subject Matter Expert (SME) – Counterterrorism Evaluation Support

Dexis

🇺🇸United States1 weeks ago
WA

Medical Expert AI Evaluation Specialist

Weekday AI

🌍Canada, United States2 weeks ago
WA

Legal AI Evaluation Expert

Weekday AI

🌍Canada, United Kingdom, United States2 weeks ago
Pantheondata logo

Flight Test and Evaluation Subject Matter Expert

Pantheondata

🇺🇸United StatesJun 5, 2026, 3:47 AM UTC