Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Toloka Ai logo

Freelance Software Engineer - AI Coding Agent Evaluation

Toloka Ai
Posted 7 hours ago
🇺🇸United States🏠Remote💰$35.0/hr📁Data & Analytics
Is this job info correct?

Please submit your CV in English and indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment. About the Role You’ll design coding tasks that challenge frontier AI coding agents. Each task is a self-contained Docker environment with a broken piece of software; an AI agent attempts the fix; automated tests verify the outcome. Your deliverable is the full task package: broken code, tests, instructions, and a reference solution proving the task is solvable. Responsibilities : Invent a realistic developer scenario — a real bug, a broken ETL, a missing feature — not a toy problem. Build a reproducible Docker environment with pinned dependencies. Write a pytest that verifies outcomes, not specific commands — deterministic, non-flaky, and does not leak the fix. Write an instruction.md that reads like a Jira ticket a developer would receive. Write a reference solve.sh proving the task is solvable. Calibrate difficulty so current state-of-the-art agents solve the task 20–60% of the time. Iterate based on feedback from expert QA reviewers. Later: review other authors’ tasks as a QA reviewer. Not in scope Data labeling, prompt engineering. Production code to ship — you design problems and verification for AI agents. Leetcode puzzles — scenarios must look like real developer work. Not every candidate task ships — quality over quantity. Requirements 3+ years of production software development in one backend stack — Python, Go, Node.js, Java, or Rust. Depth in one stack beats breadth. Python + pytest fluency — required regardless of primary stack. The task harness is pytest-based even when the broken app is in another language. Fixtures, parametrize, monkeypatch, timeouts, conftest.py. Docker authoring — reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user. Linux & Bash — comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail. AI coding agent experience — Claude Code, Cursor, Roo Code, or similar, on non-trivial work. You can cite a specific time the AI was confidently wrong and how you caught it. English — B2+ written. Not a fit Data Science, ML, or Computer Vision engineers without backend-engineering output. Manual QA testers without automation or test authoring. Frontend-only, low-code / no-code, IT Support, or Business Analysts. Engineers who have never written pytest from scratch. Junior, intern, or assistant as the most recent role. Preferred qualifications Domain depth in Security, System Administration (nginx / systemd / cron), Scientific Computing (NumPy / PyTorch / SciPy), DevOps, or Git internals. Modern Python tooling (uv, poetry, pyproject.toml). Coverage tooling (pytest-cov, coverage.py, gcov, llvm-cov, kcov). Fuzzing or property-based testing (Hypothesis). Prior contribution to agent-evaluation benchmarks or related frameworks. Process Apply → Pass qualification (90-minute sample-task screen + short behavioral interview) → Join a project → Complete tasks → Get paid. Time commitment Onboarding: ~10 hours per first task. Steady state: ~5 hours per task, 2–4 parallel tasks per author. Realistic weekly load: 8–20 hours. Higher volume available for top performers. You choose when and how to contribute; tasks must be submitted by the deadline and meet acceptance criteria. Compensation: Paid contributions, rates up to $35/hour *. Task-based compensation equivalent to hourly rate, depending on performance and volume. Some projects include incentive payments. *Rates vary based on expertise, skills assessment, location, project needs, and other factors. Higher rates may be provided to highly specialized experts. Lower rates may apply during onboarding or non-core project phases. Payment details are shared per project. Apply Submit your CV via the Mindrift platform. Indicate your English level, note this role (Software Engineering Evaluation Specialist — Terminal Bench), and include a GitHub profile link if available.

Similar jobs

Similar jobs

Toloka Ai logo

Software Engineer - AI Coding Agent Evaluation

Toloka Ai

🇺🇸United States7 hours ago
Nvidia logo

Senior Software Engineer, Agent Architecture and Evaluation

Nvidia

🇺🇸United StatesYesterday
Toloka Ai logo

AI Agent Safety Evaluation Engineer with Python - Freelance AI Trainer

Toloka Ai

🇺🇸United States3 weeks ago
CO

AI/Automation Engineer

Cohenco

🇺🇸United States16 minutes ago
Sourcing Libertymutual logo

Principal, Advanced Analytics – Internal Claims Fraud

Sourcing Libertymutual

🇺🇸United States1 hour ago
Mercor logo

AI Adversarial Specialist - Fully Remote | Upto $62/hr

Mercor

🌍Australia, Canada, Europe, United States2 hours ago