fixed-term project engagement|Site Reliability Engineer (SRE) 仕事概要 [ ABOUT TENCHIJIN ] Tenchijin Inc. is a Tokyo-headquartered space-tech company and JAXA (Japan Aerospace Exploration Agency) certified venture. Through its land-evaluation and analytics platform, Tenchijin COMPASS, the company fuses earth-observation satellite data, geospatial and environmental layers, IoT and ground-sensor feeds, and clients' operational data with proprietary AI to deliver multimodal analytics for utilities and infrastructure enterprises. Tenchijin's flagship water-infrastructure solution, KnoWaterleak, assesses pipeline deterioration and leak risk from space-derived insights and is used by a growing number of water utilities to prioritize inspection and investment, reduce non-revenue water, and extend asset life. [ ABOUT THE PROJECT ] The Global South Project is an 18-month initiative (October 2026 – March 2028) to design, build, and deploy a scalable, secure, AI-ready analytics platform tailored to utilities and infrastructure operators across Global South markets. The project adapts Tenchijin's satellite-data and GeoAI capabilities to regions facing aging or rapidly expanding infrastructure, constrained budgets, and limited field-inspection capacity — delivering risk assessment and decision-support tools that improve infrastructure management workflows. All positions are contract-based and fully remote, operating as one distributed, cross-border team with English as the working language. [ ROLE SUMMARY ] You will define what 'reliable' means for this platform and then engineer it: SLIs and SLOs for every client-facing service, proactive monitoring that catches degradation before users do, disciplined incident response, and relentless elimination of the failure modes and toil that erode reliability over time. [ Reports To ] Cloud Operations Manager / Senior DevOps Manager [ Project Language ] English (professional working proficiency or higher required) [ KEY RESPONSIBILITIES ] - Define SLIs/SLOs (availability, latency, freshness of analytics data) for platform services with the Cloud Operations Manager; implement error-budget tracking. - Build and tune monitoring, alerting, and dashboards; drive alert quality (actionable, low-noise) across the stack. - Serve in the on-call rotation; lead incident response as incident commander for major events, and author blameless postmortems with tracked follow-ups. - Engineer reliability improvements: redundancy, graceful degradation, retry/timeout policies, capacity planning, and load testing. - Automate operational toil — deployment safeguards, self-healing, runbook automation — in partnership with the Lead DevOps/SRE. - Perform production-readiness reviews for new services and tenants before launch. - Analyze reliability trends and report SLO attainment; recommend prioritization when error budgets are at risk. - Contribute chaos/failure-injection testing to validate resilience assumptions. [ WHAT WE OFFER ] - A central role in applying satellite data and AI to real infrastructure challenges, with measurable social and environmental impact in Global South markets. - Fully remote, cross-border collaboration with senior specialists across cloud, GeoAI, product, and design. - Competitive contractor compensation commensurate with experience and scope. - Direct exposure to earth-observation technology, including data ecosystems built with Japan's space agency (JAXA). Tenchijin Inc. is an equal-opportunity organization. We evaluate all applicants on qualifications and merit, without regard to nationality, race, religion, gender, age, or disability. TENCHIJIN INC. · GLOBAL SOUTH PROJECT 必須スキル [ ABOUT TENCHIJIN ] Tenchijin Inc. is a Tokyo-headquartered space-tech company and JAXA (Japan Aerospace Exploration Agency) certified venture. Through its land-evaluation and analytics platform, Tenchijin COMPASS, the company fuses earth-observation satellite data, geospatial and environmental layers, IoT and ground-sensor feeds, and clients' operational data with proprietary AI to deliver multimodal analytics for utilities and infrastructure enterprises. Tenchijin's flagship water-infrastructure solution, KnoWaterleak, assesses pipeline deterioration and leak risk from space-derived insights and is used by a growing number of water utilities to prioritize inspection and investment, reduce non-revenue water, and extend asset life. [ ABOUT THE PROJECT ] The Global South Project is an 18-month initiative (October 2026 – March 2028) to design, build, and deploy a scalable, secure, AI-ready analytics platform tailored to utilities and infrastructure operators across Global South markets. The project adapts Tenchijin's satellite-data and GeoAI capabilities to regions facing aging or rapidly expanding infrastructure, constrained budgets, and limited field-inspection capacity — delivering risk assessment and decision-support tools that improve infrastructure management workflows. All positions are contract-based and fully remote, operating as one distributed, cross-border team with English as the working language. [ ROLE SUMMARY ] You will define what 'reliable' means for this platform and then engineer it: SLIs and SLOs for every client-facing service, proactive monitoring that catches degradation before users do, disciplined incident response, and relentless elimination of the failure modes and toil that erode reliability over time. [ Reports To ] Cloud Operations Manager / Senior DevOps Manager [ Project Language ] English (professional working proficiency or higher required) [ KEY RESPONSIBILITIES ] - Define SLIs/SLOs (availability, latency, freshness of analytics data) for platform services with the Cloud Operations Manager; implement error-budget tracking. - Build and tune monitoring, alerting, and dashboards; drive alert quality (actionable, low-noise) across the stack. - Serve in the on-call rotation; lead incident response as incident commander for major events, and author blameless postmortems with tracked follow-ups. - Engineer reliability improvements: redundancy, graceful degradation, retry/timeout policies, capacity planning, and load testing. - Automate operational toil — deployment safeguards, self-healing, runbook automation — in partnership with the Lead DevOps/SRE. - Perform production-readiness reviews for new services and tenants before launch. - Analyze reliability trends and report SLO attainment; recommend prioritization when error budgets are at risk. - Contribute chaos/failure-injection testing to validate resilience assumptions. [ WHAT WE OFFER ] - A central role in applying satellite data and AI to real infrastructure challenges, with measurable social and environmental impact in Global South markets. - Fully remote, cross-border collaboration with senior specialists across cloud, GeoAI, product, and design. - Competitive contractor compensation commensurate with experience and scope. - Direct exposure to earth-observation technology, including data ecosystems built with Japan's space agency (JAXA). Tenchijin Inc. is an equal-opportunity organization. We evaluate all applicants on qualifications and merit, without regard to nationality, race, religion, gender, age, or disability. TENCHIJIN INC. · GLOBAL SOUTH PROJECT 歓迎スキル - Experience with data-platform reliability (pipeline freshness SLOs, batch-job monitoring). - Multi-region or multi-cloud reliability engineering. - Experience introducing SRE practices to teams new to them. - Experience working in a cross-functional POD / squad delivery model. 求める人物像 応募概要 給与 勤務地 Remote — Global (cross-border, distributed team) 雇用形態 Independent Contractor (renewable; engagement format at HR's discretion) 勤務体系 Contract Period: 6-month renewable contract (format at HR's discretion) — within the project window of October 2026 – March 2028 試用期間 福利厚生
Site Reliability Engineer (Senior+)
Miris
Senior Database Reliability Engineer (DBRE) (worldwide remote)
Cloudlinux
Senior Site Reliability Engineer, DevEx
Chainlink Labs
Senior Site Reliability Engineer, CCIP
Chainlink Labs
Site Reliability Platform Engineer (SRE)
Luuplicareers
Site Reliability Engineer
Supabase