TJ

Senior SRE

Moves you to
Japan
Support
Visa sponsorship
Posted
Is this job info correct?
Show job description

Our client is a leading technology company operating a large-scale digital ecosystem with a wide range of online services. The company develops and operates highly scalable technology platforms serving millions of users worldwide. With a strong focus on technology, innovation, and customer experience, the company continues to expand its digital services and invest in large-scale infrastructure and engineering capabilities.


About the Team

The Search Engineering team is responsible for developing and operating large-scale search, discovery, and navigation platforms.

The team builds highly scalable, reliable, and high-performance distributed systems that support multiple business domains and large volumes of traffic and data.

The Search Platform team is responsible for the reliability and operation of mission-critical search services. Engineers work closely with software developers, product managers, and QA engineers to continuously improve system performance, stability, and operational efficiency.


About the Position

We are looking for a highly motivated Senior Site Reliability Engineer to join the Search Engineering team.

In this role, you will be responsible for maintaining and improving the reliability of large-scale distributed systems. You will work across infrastructure, platform, and application layers, while contributing to automation, monitoring, performance optimization, and incident management.


This position is ideal for an experienced SRE or Infrastructure Engineer who enjoys solving complex technical problems and working with large-scale production environments.


Responsibilities


Your responsibilities will include, but are not limited to:


Participate in the design and deployment of production environments for new services, from infrastructure to service level


Monitor production systems and participate in on-call operations


Respond quickly to production incidents, including reporting, triage, troubleshooting, and resolution


Investigate system issues and identify root causes


Support and execute production releases, including occasional nighttime or out-of-hours operations


Improve system reliability, availability, scalability, and performance


Drive continuous improvement of operational tools and automation


Develop and maintain automation for repetitive operational tasks


Collaborate with software engineers, product managers, and QA engineers


Identify testing requirements and design appropriate testing strategies


Participate in performance testing and optimization


Support security testing and reliability exercises


Participate in disaster recovery and emergency response drills


Contribute to the improvement of engineering and operational processes


Mandatory Qualifications


- More than 8 years of work experience working in IT

- Excellent problem solving and troubleshooting skills

- Excellent teamwork and communication skills

- Strong Linux experience with understanding of system performance and reliability

- Experience analyzing and tuning performance of distributed systems

- Experience deploying and using configuration management tools (e.g., Chef, Ansible)

- Experience with observability tools (e.g., Prometheus, Grafana, Loki)

- Experience with container technologies and orchestration (Kubernetes)

- Experience automating tasks using shell scripts and/or Python

- Strong understanding of computer networking and common protocols

- Experience managing services built with Java


Desired Qualifications


- Experience automating processes for software testing and deployment (Jenkins)

- Experience working with Spark, Solr, Cassandra, or Kafka

- Experience with cloud storage solutions (Ceph, S3, MinIO)

- Experience working with a Git-based workflow

- Software development experience


Technical Environment


You may work with technologies including:


Operating Systems: Linux

Container / Orchestration: Docker, Kubernetes

Configuration Management: Chef, Ansible

Monitoring / Observability: Prometheus, Grafana, Loki

Programming / Scripting: Java, Python, Shell

CI/CD: Jenkins

Data / Distributed Systems: Spark, Solr, Cassandra, Kafka

Storage: Ceph, Amazon S3, MinIO

Version Control: Git

Infrastructure: Large-scale distributed and cloud-based environments


What We Are Looking For


We are looking for someone who:

Enjoys solving complex infrastructure and production issues

Has a strong sense of ownership for system reliability

Can investigate problems independently and identify root causes

Is comfortable working with large-scale distributed systems

Enjoys automation and continuous improvement

Can collaborate effectively with engineers and cross-functional teams

Is willing to participate in on-call and production support activities

Has a strong interest in improving system performance, stability, and scalability


Language Requirements

  • English: Fluent
  • Japanese: Optional


Work Environment

  • Flexible working hours with core collaboration hours
  • International engineering team


Salary: ¥10M – ¥13M JPY per year (Senior Level)

Location: Hybrid (4 days in the office, 1 day remote)

Visa Sponsorship: Available

Office Location: Tokyo, Japan

Working Hours: Flexible schedule with core hours from 11:00 AM to 3:00 PM


Apply now or contact us for further information:


TG Japan Inc.

03-5775-6618

RSU1_Agt@tg-hr.com

Similar jobs

Apply on LinkedIn