Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
VA

Member of Technical Staff - ML Performance

Veeda AI
Posted 6 hours ago
🌍Canada, Switzerland, United States🏢Hybrid📁Data & Analytics
Is this job info correct?

Member of Technical Staff - ML Performance About Us Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one. Responsibilities Distributed Training Throughput: Own step time and model FLOPs utilization for multi-node video world model training, choosing the tensor, context, and expert parallelism mix in PyTorch FSDP2 and Megatron-Core rather than inheriting a default. Precision & Numerical Stability: Take BF16, FP8, and NVFP4 recipes from running to converging on Blackwell, chasing scaling-factor and accumulation bugs into the video tokenizer and VAE layers where the activation outliers actually live. Kernels & Compilation: Write and tune the CUDA and Triton kernels PyTorch does not give us, driving FlashAttention-4, FlexAttention, and torch.compile integration so quadratic attention over long video sequences stops setting step time. Communication & Overlap: Tune NCCL collectives and compute/communication overlap across NVLink domains and the fabric, using the NCCL flight recorder to turn a watchdog timeout into a named rank and collective, not a restart. Fault Diagnostics & Recovery: Build the detection layer for silent data corruption (SDC), stuck CUDA kernels, and "card-freeze" hangs, plus asynchronous and tiered checkpointing that makes an interruption cost minutes rather than a day. Requirements Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field. Deep hands-on experience with PyTorch and at least one large-scale parallelism stack (FSDP2, Megatron-Core, TorchTitan, or DeepSpeed) on real multi-node jobs, not single-node approximations. Fluency in Python and C++/CUDA with the ability to predict where a kernel will stall from its memory access pattern before profiling. Experience profiling live training runs with Nsight Systems or the PyTorch profiler and translating traces into quantifiable step-time or MFU improvements. Expertise in at least one of low-precision numerics, kernel authoring, or large-run fault diagnosis, and credibility in the others. Nice to Have Experience writing Triton, CUTLASS, or CuTe-DSL kernels, or contributing to open-source kernel libraries. Experience implementing context or sequence parallelism for long-horizon video or high-token-count models. Experience running or porting large training workloads on AMD GPUs (ROCm) or Google TPUs (JAX/XLA). Experience building fault-tolerant training with elastic world size, dynamic node re-queueing, or asynchronous distributed checkpointing. Experience optimizing generative inference for interactive rollouts, including few-step samplers, distillation, and KV or latent caching. Publications or presentations on machine learning systems, compilers, or high-performance kernels. Member of Technical Staff - ML Performance About Us Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one. Responsibilities Distributed Training Throughput: Own step time and model FLOPs utilization for multi-node video world model training, choosing the tensor, context, and expert parallelism mix in PyTorch FSDP2 and Megatron-Core rather than inheriting a default. Precision & Numerical Stability: Take BF16, FP8, and NVFP4 recipes from running to converging on Blackwell, chasing scaling-factor and accumulation bugs into the video tokenizer and VAE layers where the activation outliers actually live. Kernels & Compilation: Write and tune the CUDA and Triton kernels PyTorch does not give us, driving FlashAttention-4, FlexAttention, and torch.compile integration so quadratic attention over long video sequences stops setting step time. Communication & Overlap: Tune NCCL collectives and compute/communication overlap across NVLink domains and the fabric, using the NCCL flight recorder to turn a watchdog timeout into a named rank and collective, not a restart. Fault Diagnostics & Recovery: Build the detection layer for silent data corruption (SDC), stuck CUDA kernels, and "card-freeze" hangs, plus asynchronous and tiered checkpointing that makes an interruption cost minutes rather than a day. Requirements Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field. Deep hands-on experience with PyTorch and at least one large-scale parallelism stack (FSDP2, Megatron-Core, TorchTitan, or DeepSpeed) on real multi-node jobs, not single-node approximations. Fluency in Python and C++/CUDA with the ability to predict where a kernel will stall from its memory access pattern before profiling. Experience profiling live training runs with Nsight Systems or the PyTorch profiler and translating traces into quantifiable step-time or MFU improvements. Expertise in at least one of low-precision numerics, kernel authoring, or large-run fault diagnosis, and credibility in the others. Nice to Have Experience writing Triton, CUTLASS, or CuTe-DSL kernels, or contributing to open-source kernel libraries. Experience implementing context or sequence parallelism for long-horizon video or high-token-count models. Experience running or porting large training workloads on AMD GPUs (ROCm) or Google TPUs (JAX/XLA). Experience building fault-tolerant training with elastic world size, dynamic node re-queueing, or asynchronous distributed checkpointing. Experience optimizing generative inference for interactive rollouts, including few-step samplers, distillation, and KV or latent caching. Publications or presentations on machine learning systems, compilers, or high-performance kernels.

Similar jobs

Similar jobs

Bright Vision Technologies logo

ML Performance Engineer

Bright Vision Technologies

🇺🇸United States2 days ago
GE

Senior AI/ML Performance Engineer

Generalmotors

🇺🇸United States5 days ago
Cerebras Systems logo

ML Performance Benchmarking Engineer

Cerebras Systems

🇨🇦CanadaJul 15, 2026, 4:05 PM UTC
uRun logo

Founding Engineer - ML Performance

uRun

🇺🇸United StatesJun 8, 2026, 8:03 PM UTC
RealAdvisor S.A. logo

Teamleiter Vertrieb - Coaching und Performance

RealAdvisor S.A.

🇨🇭SwitzerlandJun 6, 2026, 10:49 AM UTC
Mercor logo

Particle Physics Expert - Computational

Mercor

🌍France, Germany, India, Japan, United States33 minutes ago