Senior SRE
- Moves you to
- Japan
- Support
- Visa sponsorship
- Posted
Show job descriptionHide job description
Our client is a leading technology company operating a large-scale digital ecosystem with a wide range of online services. The company develops and operates highly scalable technology platforms serving millions of users worldwide. With a strong focus on technology, innovation, and customer experience, the company continues to expand its digital services and invest in large-scale infrastructure and engineering capabilities.
About the Team
The Search Engineering team is responsible for developing and operating large-scale search, discovery, and navigation platforms.
The team builds highly scalable, reliable, and high-performance distributed systems that support multiple business domains and large volumes of traffic and data.
The Search Platform team is responsible for the reliability and operation of mission-critical search services. Engineers work closely with software developers, product managers, and QA engineers to continuously improve system performance, stability, and operational efficiency.
About the Position
We are looking for a highly motivated Senior Site Reliability Engineer to join the Search Engineering team.
In this role, you will be responsible for maintaining and improving the reliability of large-scale distributed systems. You will work across infrastructure, platform, and application layers, while contributing to automation, monitoring, performance optimization, and incident management.
This position is ideal for an experienced SRE or Infrastructure Engineer who enjoys solving complex technical problems and working with large-scale production environments.
Responsibilities
Your responsibilities will include, but are not limited to:
Participate in the design and deployment of production environments for new services, from infrastructure to service level
Monitor production systems and participate in on-call operations
Respond quickly to production incidents, including reporting, triage, troubleshooting, and resolution
Investigate system issues and identify root causes
Support and execute production releases, including occasional nighttime or out-of-hours operations
Improve system reliability, availability, scalability, and performance
Drive continuous improvement of operational tools and automation
Develop and maintain automation for repetitive operational tasks
Collaborate with software engineers, product managers, and QA engineers
Identify testing requirements and design appropriate testing strategies
Participate in performance testing and optimization
Support security testing and reliability exercises
Participate in disaster recovery and emergency response drills
Contribute to the improvement of engineering and operational processes
Mandatory Qualifications
- More than 8 years of work experience working in IT
- Excellent problem solving and troubleshooting skills
- Excellent teamwork and communication skills
- Strong Linux experience with understanding of system performance and reliability
- Experience analyzing and tuning performance of distributed systems
- Experience deploying and using configuration management tools (e.g., Chef, Ansible)
- Experience with observability tools (e.g., Prometheus, Grafana, Loki)
- Experience with container technologies and orchestration (Kubernetes)
- Experience automating tasks using shell scripts and/or Python
- Strong understanding of computer networking and common protocols
- Experience managing services built with Java
Desired Qualifications
- Experience automating processes for software testing and deployment (Jenkins)
- Experience working with Spark, Solr, Cassandra, or Kafka
- Experience with cloud storage solutions (Ceph, S3, MinIO)
- Experience working with a Git-based workflow
- Software development experience
Technical Environment
You may work with technologies including:
Operating Systems: Linux
Container / Orchestration: Docker, Kubernetes
Configuration Management: Chef, Ansible
Monitoring / Observability: Prometheus, Grafana, Loki
Programming / Scripting: Java, Python, Shell
CI/CD: Jenkins
Data / Distributed Systems: Spark, Solr, Cassandra, Kafka
Storage: Ceph, Amazon S3, MinIO
Version Control: Git
Infrastructure: Large-scale distributed and cloud-based environments
What We Are Looking For
We are looking for someone who:
Enjoys solving complex infrastructure and production issues
Has a strong sense of ownership for system reliability
Can investigate problems independently and identify root causes
Is comfortable working with large-scale distributed systems
Enjoys automation and continuous improvement
Can collaborate effectively with engineers and cross-functional teams
Is willing to participate in on-call and production support activities
Has a strong interest in improving system performance, stability, and scalability
Language Requirements
- English: Fluent
- Japanese: Optional
Work Environment
- Flexible working hours with core collaboration hours
- International engineering team
Salary: ¥10M – ¥13M JPY per year (Senior Level)
Location: Hybrid (4 days in the office, 1 day remote)
Visa Sponsorship: Available
Office Location: Tokyo, Japan
Working Hours: Flexible schedule with core hours from 11:00 AM to 3:00 PM
Apply now or contact us for further information:
TG Japan Inc.
03-5775-6618
RSU1_Agt@tg-hr.com