Training Infrustructure Engineer
- Moves you to
- Singapore
- Support
- Relocation support
- Posted
Is this job info correct?
68,314 relocation jobs, straight from company career pages
100% free · New jobs every hour
Show job descriptionHide job description
About the Role
THE ROLE Build the training systems behind our Large Physics foundation Model. Physical data and novel architectures demand new approaches beyond the standard language-and-vision training stack. You will partner with researchers to bring experiments from prototype to scale.
Responsibilities
- Design, implement and optimise distributed training across thousands of GPUs.
- Research parallelisation and numerical precision trade-offs for novel architectures.
- Profile and debug low-level GPU operations to improve throughput and utilisation.
- Build checkpointing, fault-tolerance and reproducibility systems that hold up through rapid research iteration.
- Collaborate with researchers to scale new model architectures.
Required Skills
- Demonstrated ownership of large-model distributed training with frameworks such as FSDP, DeepSpeed, Megatron, PyTorch or JAX/XLA.
- Deep understanding of parallelism, memory optimisation, mixed precision and communication overlap.
- Ability to investigate performance from framework internals down to kernels and collectives.
- Resourcefulness, fast execution and confidence learning unfamiliar domains.
Preferred Skills
- Open-source contributions to ML infrastructure, such as PyTorch, Megatron-LM, DeepSpeed or XLA.
*This role supports relocation to Singapore.*