VA

Benjamin RLHF

Vedika API
Posted 7 hours ago
IndiaRemoteData & Analytics
Is this job info correct?

Hiring: Benjamin RL


We’re training the next generation of Vedika models.


This role sits directly inside RL and post-training research: designing how the model learns after pretraining, how it improves from interaction, how it handles long-horizon tasks, and how we push capability beyond standard instruction tuning.


You’ll work on:

• RL training for next-generation Vedika models

• GRPO, PPO, DPO and newer post-training methods

• Reward models, process rewards and verifiers

• Long-horizon reasoning and agent trajectories

• Tool-use and computer-use reinforcement

• Self-improvement and synthetic training loops

• Multi-turn behaviour and memory training

• Failure mining from model trajectories

• Evaluation systems for reasoning, autonomy and reliability

• Research experiments that can become part of the next model generation


Compensation:

₹2.6 LPA fixed

₹3.6 LPA CTC


Work mode: Fully remote


You’ll get:

• Mac for development

• Claude

• Codex

• Serious compute and research infrastructure

• ₹10L–₹50L+ yearly AI/token spend available across the team and experiments


This is not a role for someone whose idea of model work ends at prompting or basic fine-tuning.


We want someone who can understand a training run, break it, diagnose it, redesign it and make the next model measurably better.


Strong PyTorch, RL fundamentals, post-training, distributed training and hands-on experimentation matter far more than credentials.


Role: Benjamin RL

Vedika — Next Generation Models


Send your work, experiments, papers, repos or anything you trained that genuinely got better.

Similar jobs