Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Giotto.ai logo

Senior Research Engineer / Research Scientist - Post-Training, Reinforcement Learning & Training Systems

Giotto.ai
Posted 3 days ago
🇨🇭Switzerland🏠Remote📁Engineering & Development
Is this job info correct?

Giotto.ai is a Switzerland-based AI company building intelligence systems for Switzerland and Europe. Our mission is to enable governments and enterprises to retain control over the AI systems they use without compromising access to advanced reasoning capabilities. Giotto combines portable, configurable models with an AI operating system, integrating open and proprietary weights, datasets, tools, and deployment components. About the role We are looking for a Senior Research Engineer or Research Scientist to own the training and optimisation side of our complete post-training stack. Starting from pretrained checkpoints, you will design, implement, scale, and operate the methods required to produce capable, reliable, and controllable production models. Your scope will include supervised fine-tuning, preference optimisation, reinforcement learning, reward and verifier integration, policy distillation or consolidation, and distributed training. This is not a single-GPU fine-tuning or adapter-only role. You should be comfortable operating training workloads where memory, communication, rollout generation, hardware topology, and fault recovery must be designed together. You will: • Own the end-to-end post-training pipeline from pretrained checkpoint to production candidate. • Design and execute full-parameter and parameter-efficient SFT. • Implement preference optimisation, RLHF, RLAIF, reinforcement learning with verifiable rewards, and related methods. • Develop training strategies for reasoning, coding, tool use, multilingual behaviour, and long-horizon agent tasks. • Integrate reward models, verifiers, critics, graders, and process- or outcome-based rewards. • Build scalable rollout-generation systems for iterative and on-policy training. • Design multi-stage curricula combining SFT, reinforcement learning, rejection sampling, distillation, and policy consolidation. • Scale training across multiple machines and accelerators using appropriate combinations of data, tensor, pipeline, sequence, context, or expert parallelism. • Select sharding, precision, checkpointing, optimiser, batch-size, sequence-length, and activation-recomputation strategies. • Estimate memory, communication, throughput, rollout capacity, and compute requirements before launching major runs. • Profile and improve accelerator utilisation, communication efficiency, data loading, and end-to-end training time. • Diagnose numerical instability, communication failures, out-of-memory errors, stragglers, checkpoint issues, and convergence regressions. • Investigate reward hacking, entropy collapse, KL drift, stale rollouts, mode collapse, grader exploitation, and benchmark overfitting. • Build reliable checkpointing, recovery, monitoring, and reproducibility procedures. • Collaborate closely with data, evaluation, infrastructure, and inference teams. • Contribute clean, tested code, technical reports, and operational runbooks. We are looking for demonstrated experience in most of the following areas: • Ownership of large-scale language-model training or post-training runs across multiple machines and accelerators. • Experience with workloads for which straightforward single-node training or pure data parallelism was insufficient.. • Deep proficiency with Python, PyTorch, autograd, mixed precision, optimisation, and distributed execution. • Practical experience with PyTorch Distributed, FSDP, DeepSpeed, Megatron-Core, or an equivalent framework. • Ability to select parallelism and sharding strategies based on model, sequence, memory, and network constraints. • Strong understanding of SFT, preference optimisation, reinforcement learning, reward modelling, KL regularisation, sampling, and training stability. • Experience operating high-throughput inference or rollout systems as part of a training loop. • Ability to debug across model code, distributed communication, numerical optimisation, data, and infrastructure. • Strong experimental design and the ability to distinguish algorithmic improvements from evaluation or systems artefacts. • Experience building reliable, observable, and reproducible research software. • Personal ownership of consequential decisions affecting a substantial training programme. A PhD is not required. We value exceptional technical work, strong judgement, and demonstrated ownership. Relevant stack • Python and PyTorch. • PyTorch Distributed and FSDP. • DeepSpeed, Megatron-Core, or comparable frameworks. • Hugging Face Transformers. • CUDA and NCCL. • vLLM, SGLang, or similar rollout engines. • Ray, Slurm, Kubernetes, or comparable orchestration systems. • MLflow or Weights & Biases. • Docker, GCP, GitLab CI, profiling, monitoring, and pytest. Experience with CUDA or Triton, long-context training, sparse models, asynchronous RL, stateful agent environments, distillation, or deployment-aware post-training would be especially valuable. You may be a strong fit if you: • Enjoy working at the intersection of model research and distributed systems. • Can move from paper reproduction to reliable scaled implementation. • Are comfortable taking responsibility for expensive and operationally demanding experiments. • Approach failures methodically across algorithms, data, numerical stability, and infrastructure. • Care about held-out capability and reliability, not only training loss or reward. • Want meaningful ownership of a complete model programme. Location and work style We offer full-time employment in Switzerland. • Remote work is supported. • The team gathers approximately one week per month in a Swiss office. • Exceptional candidates elsewhere in Europe may be considered..

Similar jobs

Similar jobs

SP

.NET Developer_BIOGGIO_Lugano (Svizzera)

Spindox

🇨🇭Switzerland3 hours ago
BDO Switzerland logo

IT Project Leader (w/m/d) 60-100%

BDO Switzerland

🇨🇭Switzerland3 hours ago
PK

IT System Engineer (100%)

Pallas Kliniken

🇨🇭Switzerland3 hours ago
BR

Executive Director, Medical Communications & Digital Excellence

Bristolmyerssquibb

🌍Switzerland, United Kingdom, United States3 hours ago
LocalBini logo

Growth Digital Marketing Specialist

LocalBini

🇨🇭Switzerland3 hours ago
PraxTor logo

Tarifspezialist ambulante Tarife 80-100% (m/w)

PraxTor

🇨🇭Switzerland3 hours ago