Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
VA

Member of Technical Staff - ML Operations

Veeda AI
Posted 6 hours ago
🌍Canada, Switzerland, United States🏢Hybrid📁Data & Analytics
Is this job info correct?

ABOUT US Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one. RESPONSIBILITIES Run Lifecycle & Launch Tooling: Own how a run is defined, launched, resumed, and killed, from typed configs (Hydra, OmegaConf) to pinned container digests to a relaunch that takes one command. Experiment Tracking & Provenance: Bind every checkpoint to its code commit, config hash, dataset version, and container digest in Weights & Biases or MLflow, so an old run rebuilds from its manifest, not from memory. Checkpoint Registry & Lineage: Own retention and garbage-collection policy, PyTorch DCP resharding and format conversion, and promotion from raw checkpoint to evaluated artifact that simulation and robotics can safely build on. Evaluation in CI: Gate each checkpoint on seeded rollout and policy-success suites, run per-change and nightly as Slurm arrays under Argo Workflows, with confidence intervals wide enough to separate regression from eval noise. Goodput & Incident Response: Carry the pager for live runs (loss spikes, throughput cliffs, data loader stalls), and report goodput against allocated GPU-hours as the number capacity decisions actually run on. REQUIREMENTS You have a Bachelor's degree or equivalent hands-on experience in Computer Science, Engineering, or a related technical field. You have strong Python and software engineering skills and real CI/CD experience (e.g., GitHub Actions, Buildkite), and you have shipped internal tooling that other engineers chose to keep using. You have operated multi-node training jobs, carried the pager for them, and decided from telemetry whether to kill, requeue, or let a degraded run ride. You have built reproducible pipelines end to end, and can say precisely which parts of a training run are bit-reproducible, which are not, and why. You are rigorous about evaluation methodology, from seeds and sample sizes to confidence intervals, and can tell a genuine regression from a flaky harness. NICE TO HAVE You have run experiment tracking at scale, logging video, 3D, and trajectory artifacts rather than only scalars. You have built evaluation harnesses for generative or embodied models, where quality is a distribution rather than a pass/fail. You have orchestrated ML workflows with Argo Workflows, Flyte, or Ray, and know where each one breaks. You are fluent with Prometheus, Grafana, and OpenTelemetry, and instrument a training job before its first outage. You have built GPU-hour attribution that maps cluster spend back to specific experiments and teams. You have written a postmortem that changed how a team ran jobs, not just what it logged. You have contributed to open-source ML tooling, or published on evaluation or reproducibility methodology.

Similar jobs

Similar jobs

Pathward, N.A. logo

Manager, Sanctions Program — BSA/AML & Sanctions Operations

Pathward, N.A.

🇺🇸United StatesYesterday
Pathward, N.A. logo

AVP, Program Governance — BSA/AML & Sanctions Operations

Pathward, N.A.

🇺🇸United StatesYesterday
Wise logo

FinCrime Operations Team Lead - AML Investigations

Wise

🇺🇸United States1 weeks ago
Fabley logo

AI/ML Operations Engineer (MLOps Engineer)

Fabley

🇨🇭Switzerland1 weeks ago
Revenue Pulse logo

Marketing Operations Specialist (HTML / Liquid)

Revenue Pulse

🇨🇦CanadaJul 18, 2026, 9:06 AM UTC
Aaaie logo

Software Engineer III - AI/ML Platform Operations - Remote

Aaaie

🇺🇸United StatesJul 2, 2026, 8:43 AM UTC