ABOUT US Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one. RESPONSIBILITIES Run Lifecycle & Launch Tooling: Own how a run is defined, launched, resumed, and killed, from typed configs (Hydra, OmegaConf) to pinned container digests to a relaunch that takes one command. Experiment Tracking & Provenance: Bind every checkpoint to its code commit, config hash, dataset version, and container digest in Weights & Biases or MLflow, so an old run rebuilds from its manifest, not from memory. Checkpoint Registry & Lineage: Own retention and garbage-collection policy, PyTorch DCP resharding and format conversion, and promotion from raw checkpoint to evaluated artifact that simulation and robotics can safely build on. Evaluation in CI: Gate each checkpoint on seeded rollout and policy-success suites, run per-change and nightly as Slurm arrays under Argo Workflows, with confidence intervals wide enough to separate regression from eval noise. Goodput & Incident Response: Carry the pager for live runs (loss spikes, throughput cliffs, data loader stalls), and report goodput against allocated GPU-hours as the number capacity decisions actually run on. REQUIREMENTS You have a Bachelor's degree or equivalent hands-on experience in Computer Science, Engineering, or a related technical field. You have strong Python and software engineering skills and real CI/CD experience (e.g., GitHub Actions, Buildkite), and you have shipped internal tooling that other engineers chose to keep using. You have operated multi-node training jobs, carried the pager for them, and decided from telemetry whether to kill, requeue, or let a degraded run ride. You have built reproducible pipelines end to end, and can say precisely which parts of a training run are bit-reproducible, which are not, and why. You are rigorous about evaluation methodology, from seeds and sample sizes to confidence intervals, and can tell a genuine regression from a flaky harness. NICE TO HAVE You have run experiment tracking at scale, logging video, 3D, and trajectory artifacts rather than only scalars. You have built evaluation harnesses for generative or embodied models, where quality is a distribution rather than a pass/fail. You have orchestrated ML workflows with Argo Workflows, Flyte, or Ray, and know where each one breaks. You are fluent with Prometheus, Grafana, and OpenTelemetry, and instrument a training job before its first outage. You have built GPU-hour attribution that maps cluster spend back to specific experiments and teams. You have written a postmortem that changed how a team ran jobs, not just what it logged. You have contributed to open-source ML tooling, or published on evaluation or reproducibility methodology.
Manager, Sanctions Program — BSA/AML & Sanctions Operations
Pathward, N.A.
AVP, Program Governance — BSA/AML & Sanctions Operations
Pathward, N.A.
FinCrime Operations Team Lead - AML Investigations
Wise
AI/ML Operations Engineer (MLOps Engineer)
Fabley
Marketing Operations Specialist (HTML / Liquid)
Revenue Pulse
Software Engineer III - AI/ML Platform Operations - Remote
Aaaie