Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Comfy logo

Senior/Staff AI Cloud Infra Engineer

Comfy
Posted May 27, 2026, 9:32 PM UTC
🇺🇸United States🏠Remote📁Engineering & Development
Is this job info correct?

The Role: We're looking for a Cloud Infrastructure Engineer who thrives on building and scaling large-scale GPU compute platforms fast. You'll be instrumental in developing and managing the foundational infrastructure that powers our AI workloads. Our core infrastructure relies heavily on Python, Kubernetes (K8s), Terraform, and Ansible, but we care more about your ability to learn, adapt, and ship robust solutions than whether you've used these exact tools before. You are a good fit if this describes you: You excel at building and managing distributed compute platforms, especially those involving GPU resources. You have deep expertise in backend systems that orchestrate complex workloads efficiently, managing capacity and resource constraints. You possess a strong understanding of foundational cloud infrastructure (AWS/GCP/Azure) and Linux provisioning/management tools. You know how to design for reliability and scale with minimal operational overhead. You learn new technologies rapidly because you're excited by solving hard infrastructure challenges. You've scaled infrastructure before and understand the tradeoffs that matter. You think most infrastructure moves too slowly and could be way better automated and optimized. You're comfortable diving into unfamiliar systems and making them work reliably. You are a self-starter who executes quickly, takes ownership, and constantly seeks improvement. What you'll do: Develop and maintain our core Python platform for routing requests, orchestrating AI workloads, managing GPU server capacity, observability, and more. Develop and maintain our infrastructure layer using Terraform, Ansible, and cloud provider APIs to manage our fleet of GPU workers across cloud and potentially bare metal environments. Own and operate the technologies underpinning our platform, potentially including K8s, FluxCD, Nomad, Prometheus, Thanos, Grafana, Loki, distributed networking/storage, etc. Architect and implement solutions that directly impact the performance and availability of services for millions of ComfyUI users. Work closely with our core engineering team to design and build new infrastructure systems. Help create the vision and lay the foundation for where our infrastructure should go in the next 1/2/5 years. Help shape our technical direction and infrastructure best practices as we grow. Requirements: Deep experience building and managing distributed compute platforms, preferably using Python. Strong foundation in managing cloud infrastructure (AWS, GCP, or Azure). Experience with bare metal is a plus. Solid understanding of container orchestration (Kubernetes preferred) and CI/CD principles and tools. Excellent communication skills. Proven ability to learn fast and ship quality infrastructure code and configurations. Nice to have: You have excelled at a fast-paced, high-growth tech startup before or are extremely excited about being in one. Experience specifically with GPU management, scheduling, and monitoring in a large-scale environment. Experience with specific observability tools (Prometheus, Grafana, Loki, Thanos).

Similar jobs

Similar jobs

GS

Software Engineer, Infrastructure

Gray Swan AI

🇺🇸United States8 hours ago
AA

Infrastructure Engineer

Applied Atomics

🇺🇸United States8 hours ago
Brook & Whittle logo

Senior Infrastructure Engineer

Brook & Whittle

🇺🇸United States14 hours ago
Bold New Solutions - BNS Power logo

Sales Engineer Data Center & AI Infrastructure

Bold New Solutions - BNS Power

🇺🇸United States15 hours ago
Nvidia logo

Senior Software Engineer, Infrastructure Automation and Distributed Systems

Nvidia

🇺🇸United States15 hours ago
Nvidia logo

Senior Deep Learning Infrastructure Engineer

Nvidia

🇺🇸United States15 hours ago