Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Odixcity Consulting logo

RLHF Specialist

Odixcity Consulting
Posted 2 days ago
🌍Kenya, Nigeria, South Africa, Spain🏠Remote📁Data & Analytics
Is this job info correct?
Job Title: RLHF Specialist

Location: Remote (Worldwide)

Job Summary: An RLHF Specialist is responsible for improving and aligning AI models using Reinforcement Learning from Human Feedback (RLHF) methodologies. This role focuses on designing, implementing, and optimizing feedback pipelines that enhance model performance, safety, factual accuracy, and alignment with human values.

Responsibilities

  • Generate high-quality preference data by comparing multiple model responses and ranking them based on criteria such as helpfulness, honesty, and harmlessness (HHH).
  • Design complex, multi-turn prompts to stress-test model behavior and expose weaknesses in reasoning or safety.
  • Write detailed “chain-of-thought” explanations and rationales to train reward models on why specific responses are superior.
  • Collaborate with Machine Learning Engineers to analyze model failure modes and identify data gaps that, when filled, will improve reinforcement learning outcomes.
  • Develop and iterate on annotation strategies for preference scoring and reinforcement signals, ensuring consistency across a global team.
  • Proactively probe models to identify vulnerabilities, biases, or hallucination patterns, documenting findings for model optimization.
  • Analyze edge cases where the reward model behaves unexpectedly (e.g., over-indexing on verbosity or style over substance). Provide detailed feedback to ML engineers on reward model failure modes and suggest specific data interventions to correct model behavior.
  • Develop and document templated instruction sets for larger annotation teams. Translate complex reinforcement learning concepts into simple, repeatable tasks for junior reviewers, ensuring high-quality data collection at scale.
  • Monitor model performance over time by maintaining a personal test set of prompts. Regularly re-evaluate new model versions against historical benchmarks to track improvements or regressions in reasoning and alignment.

Requirements

  • Minimum of 2 years of experience in Data Annotation, Model Evaluation, Computational Linguistics, or Trust and Safety, specifically working with AI/ML training data.
  • Strong proficiency in Python and deep learning frameworks (PyTorch, JAX, or TensorFlow).
  • Deep understanding of Reinforcement Learning concepts (PPO, Trust Regions, Reward Hacking) and how they apply to language generation.
  • Hands-on experience fine-tuning open-source models (e.g., Llama 2/3, Mistral, gemma) using techniques like LoRA/QLoRA.
  • Experience working with annotation tools (LabelBox, Scale AI, Snorkel) and managing human-in-the-loop workflows.
  • Ability to diagnose why an RL policy collapsed and adjust hyperparameters or reward structure accordingly.
  • Experience with Constitutional AI or Self-Alignment techniques.
  • Contributions to open-source alignment libraries (TRL, Transformer Reinforcement Learning, Axolotl).
  • Experience with cloud Platforms (AWS SageMaker, GCP Vertex AI).

Similar jobs

Similar jobs

Automatit logo

Senior Data Architect

Automatit

🌍Israel, Spain, United Kingdom3 hours ago
Medtronic logo

Senior Product Specialist Robotic Surgical Technologies Iberia

Medtronic

🇪🇸Spain2 hours ago
Gridlines logo

Manager - Financial Model Build

Gridlines

🇿🇦South Africa2 hours ago
Cnx logo

Fashion Consultant (German-speaking) - Remote - Sport Clothing Industry PM01

Cnx

🇪🇸Spain2 hours ago
African Safari Group logo

Business Technology Architect

African Safari Group

🇿🇦South Africa2 hours ago
African Safari Group logo

Head of Sales

African Safari Group

🇿🇦South Africa2 hours ago