Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Gauss Labs logo

Senior Site Reliability Engineer (KR)

Gauss Labs
Posted 1 weeks ago
🇰🇷South Korea🏢Hybrid📁Engineering & Development
Is this job info correct?

Gauss Labs is an industrial AI company on a mission to revolutionize manufacturing with AI, starting with the semiconductor sector. Panoptes is an AI-based virtual metrology solution deployed in high-volume manufacturing fabs, helping customers improve yield, reduce costs, and accelerate production. Our software runs in our customers' own managed environments, and we're seeking a Site Reliability Engineer to own the reliability of the infrastructure and platform that Panoptes runs on. You will keep the platform available, performant, and scalable; own monitoring, alerting, incident first-response, and the on-call rotation; and build the automation and observability that let engineering teams operate their services safely. Responsibilities Platform reliability and operations: Own platform-layer reliability across both environments. In our internal cloud environment: full ownership — cluster health, resource management (CPU/memory/OOM), scheduling, autoscaling, Kubernetes/EKS lifecycle. In the customer environment: operate directly at the application-namespace level and for the customer-controlled cluster/node layer, diagnose and clearly communicate what's needed, and operate the platform within their setup, decisions, and constraints. Monitoring and Alerting: Build and maintain robust monitoring and alerting for the infrastructure and platform layer to proactively identify and resolve issues before they impact the platform. Incident Response: Own incident first-response for the platform layer and participate in the on-call rotation to minimize downtime and restore service quickly. Automation: Develop automation tools and scripts to streamline operations, reduce manual effort, and enable engineering teams to operate their own services safely. Capacity Planning: Forecast resource needs, optimize resource utilization, and ensure the platform infrastructure can handle increasing workloads. Deployment infrastructure: Build and maintain CI/CD pipelines and deployment infrastructure for the platform. Continuous Improvement: Drive a culture of continuous improvement by identifying opportunities to enhance platform reliability, performance, and efficiency. Basic Qualifications Bachelor's degree in computer science, engineering, or a related discipline 5+ years of industry experience as a Site Reliability Engineer or in platform/infrastructure engineering Hands-on experience operating Kubernetes in production (EKS preferred): cluster lifecycle, scheduling, autoscaling, resource management Experience with cloud platforms (AWS preferred) and containerization technologies (Docker, Kubernetes) Experience with observability and alerting tools (Prometheus, Grafana, ElasticSearch, Jaeger) Experience with scripting languages (Python, Bash) Working knowledge of GitHub, GitHub Actions, and CI/CD concepts Strong problem-solving and troubleshooting skills Working proficiency in English for internal documentation and technical coordination Preferred Qualifications Knowledge of AI/ML infrastructure and workloads. Knowledge of database technologies (MongoDB, PostgreSQL) Experience operating software in customer-managed (on-prem or customer-cloud) environments Exposure to manufacturing, semiconductor, or enterprise B2B customer environments [Interview process] Application reivew - Phone interview - Virtual onsite interview - VP interview/Core Value interview - CEO interview

Similar jobs

Similar jobs

MinIO logo

Site Reliability Engineer - South Korea

MinIO

🇰🇷South Korea2 weeks ago
Furiosa Ai logo

Software Engineer, Site Reliability Engineer

Furiosa Ai

🇰🇷South KoreaMay 29, 2026, 12:24 AM UTC
Mercor logo

Electrical Engineer - Fully Remote | Upto $100/hr

Mercor

🌍India, South KoreaYesterday
Darktrace logo

Cyber Support Engineer

Darktrace

🇰🇷South KoreaYesterday
Salesforce logo

Technical Architect

Salesforce

🌍Mexico, Philippines, South Korea2 days ago
BE

Senior System Designer, Overwatch

Blizzard Entertainment

🇰🇷South Korea16 hours ago