Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
WA

QA/Test Engineer

Weekday AI (YC W21)
Posted 2 days ago
🇺🇸United States🏠Remote💰$60.0–$90.0/hr📁Engineering & Development
Is this job info correct?
This role is for one of our clients

Compensation: $60-$90 per hour

Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced QA and Test Engineers to ensure every benchmark is reliable, reproducible, and accurately measures real AI capabilities.

In this role, you will review complex, multi-step evaluation tasks, validate their correctness, identify edge cases, and strengthen testing methodologies before benchmarks are deployed. You'll collaborate closely with AI researchers and task authors to improve evaluation quality, eliminate ambiguity, and ensure benchmark integrity.

This is a fully remote, full-time engagement requiring approximately 35 hours per week.

Requirements

Key Responsibilities

  • Design comprehensive test cases that validate evaluation tasks, including complex edge cases and unexpected scenarios
  • Review benchmark tasks and reference solutions to identify ambiguity, inconsistencies, missing requirements, and grading gaps
  • Debug task environments and Python-based evaluation scripts to ensure reliable execution and accurate results
  • Develop repeatable quality assurance processes, validation checklists, and testing frameworks for benchmark creation
  • Identify potential shortcuts, exploits, or weaknesses that could compromise evaluation accuracy or benchmark integrity
  • Collaborate with AI researchers, engineers, and task authors to improve task quality, reproducibility, and technical rigor

Required Qualifications

  • Master's degree, PhD, or equivalent practical experience in a STEM discipline involving software engineering, research, or advanced technical problem solving
  • Minimum 1 year of professional experience in Quality Assurance, Test Engineering, Software Engineering, Research Engineering, or a related technical field with strong quality ownership
  • Proven experience designing test cases, validating complex software systems, and debugging end-to-end workflows
  • Strong proficiency in Python and Git, with the ability to troubleshoot unfamiliar codebases and technical environments
  • Excellent analytical thinking, problem-solving skills, and exceptional attention to detail
  • Experience documenting bugs, test strategies, and technical findings with clear written communication
  • Experience evaluating AI systems, machine learning models, or AI-generated outputs is preferred
  • Ability to work independently while managing multiple complex tasks with minimal supervision
  • Ability to commit approximately 35 hours per week on a consistent basis

Preferred Qualifications

  • Experience with AI evaluation, benchmark development, or quality assurance for machine learning systems
  • Background in automation testing, validation frameworks, or software quality engineering
  • Familiarity with large language models, AI agent workflows, or evaluation pipelines
  • Experience creating repeatable QA processes for research or engineering projects

Why Join

  • Play a key role in improving the quality and reliability of next-generation AI evaluation benchmarks
  • Collaborate with leading AI researchers on cutting-edge evaluation methodologies
  • Help ensure AI systems are tested against rigorous, real-world scenarios
  • Apply your software testing and quality engineering expertise to advance AI reliability
  • Enjoy the flexibility of a fully remote engagement while contributing to impactful AI research

Equal Opportunity

We are committed to creating an inclusive workplace where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.

Contract & Engagement Details

  • Independent contractor engagement
  • Fully remote with flexible working hours
  • Expected commitment of approximately 35 hours per week
  • Project duration may be extended, shortened, or concluded based on project requirements and individual performance
  • Work does not require access to confidential or proprietary information from any current or former employer
  • Payments are issued weekly based on approved work completed
  • At this time, we are unable to support H1-B or STEM OPT candidates

Similar jobs

Similar jobs

DD

Sr. Software Engineer in Test

DDN

🇺🇸United States7 hours ago
DD

Staff Engineer in Test

DDN

🇺🇸United States7 hours ago
Nokia logo

Staff PIC RF Test Development Engineer

Nokia

🇺🇸United States14 hours ago
Muon Space logo

Senior Software Engineer, Hardware Test

Muon Space

🇺🇸United States14 hours ago
DD

Senior Enterprise Systems Test Engineer

Ddcdine

🇺🇸United States15 hours ago
DD

Automated Test Engineer

Ddcdine

🇺🇸United States15 hours ago