Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
FriendliAI logo

Software Engineer – Cloud Infrastructure

FriendliAI
Posted 3 hours ago
🇺🇸United States🏢Hybrid📁Engineering & Development
Is this job info correct?

About the job FriendliAI is looking for a Cloud Infrastructure Engineer to own the architecture and evolution of the cluster platform behind our GPU-accelerated AI inference cloud. As a Software Engineer, Cloud Infrastructure, you will design how our clusters are built and connected, extend Kubernetes where its defaults fall short, and own the network path that inference traffic depends on. Inference is an unforgiving workload for Kubernetes. Traffic is bursty and latency-sensitive, GPU capacity is scarce and inelastic, tenants must stay isolated, and multi-node serving depends on the network holding up under sustained load. This is a hands-on architecture role for an engineer who has already run large clusters in production and wants to push them further. Key Responsibilities Cluster Architecture Own the architecture of our multi-cluster, multi-tenant Kubernetes fleet across both managed and self-managed clusters: cluster topology, control plane and etcd lifecycle, and zero-downtime upgrades. Extend Kubernetes with custom controllers, operators, and CRDs so platform behavior is encoded in software rather than runbooks. Design GPU scheduling and capacity strategy, including topology-aware placement, node pools, priority and preemption, and quota across tenants. Build autoscaling that matches inference traffic: queue-driven pod scaling, node autoscaling, scale-to-zero, and cold-start reduction. Networking Own the Kubernetes network data plane: CNI, IPAM, DNS, ingress, and L4/L7 load balancing. Design cross-AZ, cross-region, and cross-cluster connectivity, and operate the service mesh for routing, mTLS, and traffic policy. Debug production network issues (packet loss, conntrack exhaustion, MTU mismatches, DNS latency, load balancer behavior) and drive permanent fixes. Reliability & Collaboration Define SLOs for platform-critical systems and lead post-incident hardening. Deliver infrastructure as code with Terraform, Helm, and GitOps. Partner with the inference engine, platform, SRE, and security teams to turn serving requirements into platform capabilities. Qualifications 5+ years designing, building, and operating large-scale Kubernetes infrastructure in production. Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent. Proven experience operating large-scale, high-traffic network services in production. Deep understanding of Kubernetes internals: API server, scheduler, controller loops, kubelet, and etcd. Strong command of Kubernetes and cloud networking: CNI, kube-proxy/eBPF datapaths, DNS, load balancing, service mesh, and VPC routing. Proficiency with AWS, Terraform, Helm, and Ansible. Programming skills in Go or Python, with the ability to build infrastructure tooling and automation. Strong debugging skills across distributed systems, containers, and the Linux networking stack. Clear written and verbal communication, including the ability to document architectural decisions for other engineers. Preferred Experience Large-scale Kubernetes operations in a high-traffic domain such as gaming, e-commerce, or public cloud. Cilium and eBPF, including kube-proxy replacement or upstream contributions. Cluster provisioning and lifecycle management with Kubespray or similar Ansible-based tooling. GPU orchestration: NVIDIA GPU Operator, device plugins, or Dynamic Resource Allocation (DRA). High-performance networking for distributed workloads: RDMA/RoCE, InfiniBand, EFA, SR-IOV, or NCCL tuning. Multi-cloud, hybrid-cloud, or bare-metal Kubernetes operations. Contributions to Kubernetes, Cilium, Istio, or other CNCF projects. Benefits Flexible working hours Daily lunch and dinner provided; unlimited snacks and beverages Supportive and highly collaborative work environment Health check-up support and top-tier equipment/hardware support A front-row seat to the generative AI infrastructure revolution Competitive compensation, startup equity, health insurance, and other benefits. About FriendliAI FriendliAI is the fastest inference cloud for agents, built to run frontier open-weight models in production at scale. It delivers up to 7x faster output token speed, up to 90% lower inference costs, and 99.99% uptime across the most demanding agent workloads — long-context inference, real-time streaming, and accurate tool calling. We are a small, fast-moving team doing work that matters at one of the most exciting moments in the history of technology. With our world-class inference stack, we are building the platform teams can actually rely on.

Similar jobs

Similar jobs

Hcsc logo

PaaS Infrastructure Engineering Cloud Automation &Tools Consultant - WFH

Hcsc

🇺🇸United States12 hours ago
Huron logo

Cloud Infrastructure + Security Architect

Huron

🇺🇸United States2 days ago
Bright Vision Technologies logo

Cloud Infrastructure Engineer – AWS

Bright Vision Technologies

🇺🇸United States2 days ago
Visualedgeit logo

IT Systems Engineer – Microsoft & Cloud Infrastructure

Visualedgeit

🇺🇸United States2 days ago
Knitwellgroup logo

Cloud Infrastructure Engineer (Multi-Cloud/Terraform)

Knitwellgroup

🇺🇸United States2 days ago
Empower Pharmacy logo

Staff Cloud Infrastructure Architect

Empower Pharmacy

🇺🇸United States4 days ago