Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Cox Exponential logo

Founding Engineer, AI Infra

Cox Exponential
Posted Jun 15, 2026, 7:59 PM UTC
🇺🇸United States🏢Hybrid📁Engineering & Development
Is this job info correct?

About Goaly At Goaly, our mission is to make custom AI affordable for every business. Our founding team comes from the front lines of top AI labs and tech giants (Meta MSL, TikTok AI, Google DeepMind, xAI, Microsoft Research, etc.), where we built large-scale training infrastructure powering trillion-parameter models and scaled GenAI models to a global user base. Now, we are building something we wish we had before: a platform that makes training and adapting custom AI affordable for all modern companies, not just Big Tech. Our north star is ambitious: for a domain-specific task, reach 90% of SOTA performance at less than 10% of the cost. To get a taste of what we are doing, see our first tech blog . About the Role You will sit at the intersection of systems engineering and applied ML, building specialized infrastructure that keeps large language and multimodal models fast, reliable, and cost-effective. You will partner with research, product, and infra teams to ship production-ready platforms for training and serving AI at scale. Key Responsibilities Efficiency & performance : Improve LLM training and inference efficiency through better memory utilization, optimized parallelism, and kernel-level innovations (e.g. FlashAttention, CUDA/Triton). Training & RL robustness : Build scalable, stable training and RL pipelines with strong reproducibility, observability, and debuggability. Serving & inference optimization : Design and tune high-throughput, low-latency model serving systems, including quantization, caching, and speculative decoding. Scalability & infrastructure : Own end-to-end training and inference infrastructure — from data ingestion and checkpointing to multi-GPU and multi-cloud orchestration. Production enablement : Work closely with researchers and product engineers to turn new algorithms into reliable, production-ready systems. Requirements 5+ years building or operating ML infrastructure at scale, ideally supporting large language or multimodal models. Deep understanding of GPU architecture, distributed training frameworks (PyTorch, DeepSpeed, Megatron, Ray), and parallelism strategies. Hands-on experience running inference stacks (vLLM / SGLang, TGI, Triton) and optimizing them via low-level profiling. Strong software engineering fundamentals in Python and one of C++/Rust/Go, with clean, reliable code shipped to production. Working knowledge of modern data pipelines, feature stores, and vector databases used in production AI systems. Comfort automating infrastructure with Kubernetes, Terraform/Pulumi, and observability stacks (Prometheus, Grafana, OpenTelemetry). Bonus Points Experience deploying open-source LLMs (Llama 3, Qwen, DeepSeek) or training custom foundation models. Contributions to ML systems tooling (compilers, kernels, inference runtimes) or open-source infrastructure projects. Background in reinforcement learning, evaluation harnesses, or alignment tooling that hardens production AI systems.

Similar jobs

Similar jobs

Bold New Solutions - BNS Power logo

Business Development Manager Data Center & AI Infrastructure

Bold New Solutions - BNS Power

🇺🇸United States8 hours ago
Bold New Solutions - BNS Power logo

Sales Engineer Data Center & AI Infrastructure

Bold New Solutions - BNS Power

🇺🇸United States8 hours ago
Usbank logo

Lead Software Engineer (Hogan Mainframe)

Usbank

🇺🇸United States8 hours ago
DD

Mainframe Solutions Architect

Ddcdine

🇺🇸United States8 hours ago
DD

Mainframe Ops/App Support

Ddcdine

🇺🇸United States8 hours ago
Navy Federal Credit Union logo

Principal Mainframe Engineer

Navy Federal Credit Union

🇺🇸United States9 hours ago