Platform Reliability Engineer Reports To: Platform Reliability Engineering Manager Location: Remote (within the U.S.) Environment: Remote Status: Exempt; Salaried Who We Are: Recognized by Gartner in their Modern 4PL Market Guide, Redwood Logistics is at the forefront of industry innovation. Our cutting-edge supply chain technology pairs with the expertise of our brilliant minds to empower logistics execution across North America and Mexico. Leveraging a comprehensive range of services, data-centric network solutions, and a seamlessly integrated platform, we have established our prominence as a key player in the mid-market segment within the freight tech industry. Whether you’re just starting your career or are an established professional looking for your next opportunity, Redwood inspires innovation across teams to provide transformative solutions for our customers. Purpose of Your Work: As a Platform Reliability Engineer , you will design, build, and evolve Redwood's internal engineering platform, enabling product teams to deliver secure, scalable, and highly reliable software. You will create shared cloud infrastructure, Kubernetes platforms, infrastructure automation, CI/CD, observability, and developer self-service capabilities while driving operational excellence across the engineering organization. Working closely with Software Engineering, Architecture, Security, Data Engineering, and MLOps, you will establish engineering standards, improve platform reliability, eliminate operational toil through automation, and enable developers to deliver software faster with confidence. How You Make a Difference Everyday: Design and maintain reusable Terraform modules, Helm charts, GitOps templates, and shared platform services. Develop self-service platform capabilities and golden paths that accelerate software delivery. Design, standardize, and evolve enterprise CI/CD pipelines incorporating automated testing, security validation, deployment governance, and release automation. Engineer highly available Kubernetes platforms with strong security, networking, scalability, and operational practices. Drive platform observability through metrics, logs, traces, dashboards, alerting, SLIs, SLOs, and error budgets. Support engineering response during production incidents, perform post-incident reviews, and implement permanent reliability improvements. Continuously eliminate manual operational work through automation. Partner with engineering teams to improve application reliability, resilience, performance, and operational readiness. Optimize cloud cost, platform efficiency, and resource utilization. You’ve Got This? 5+ years in Platform Engineering, Site Reliability Engineering, or Cloud Infrastructure. Production experience with Azure and/or AWS. Production Kubernetes experience (AKS, EKS, or GKE). Infrastructure as Code using Terraform or equivalent. Modern CI/CD platforms (GitHub Actions, Azure DevOps, or equivalent). GitOps using ArgoCD or Flux. Helm chart development. Linux systems administration and container technologies. Scripting using PowerShell, Bash, Python, or Go. Enterprise observability platforms such as Datadog, Prometheus, Grafana, LogicMonitor, or New Relic. Cloud networking and identity technologies including OIDC. Strong troubleshooting, communication, and cross-functional collaboration. Strong communication and collaboration skills with experience partnering across Software Engineering, Cloud Infrastructure, Architecture, Security, Data Engineering, and MLOps teams. What We Offer: Access to experts and resources for your Learning & Development journey Opportunity for internal mobility Employee referral bonus program Employee Resource Groups (ERGs) Annual fundraising and volunteer events to give back to communities Paid time off, floating holidays, time off to volunteer and rollover Paid parental leave Medical, dental, vision and 401k plans (with match) Flexible spending account, mass transit and dependent care plans available Health savings account, with a annual company contribution for plan participants Short-term and long-term disability; life insurance policies subsidized by company Additional benefits including pet insurance, accident care, access to legal advice and more Work Schedule: This position is full-time and remote Monday through Friday from 8:00 AM to 5:00 PM with an hour break, but flexibility is available based on coverage. Compensation Range: Salary Range: $160,000 - $175,000 This position is eligible to earn annual incentives based on individual and company performance. The estimated pay range reflects an anticipated range for this position. The actual base salary offered will depend on a variety of factors, including the qualifications of the individual applicant for the position, years of relevant experience, specific and unique skills, level of education attained, certifications or other professional licenses held, and the geographical location in which the applicant lives and/or which they will be performing the job. Redwood is an equal opportunity employer. Employment decisions at the Company are based on individual merit, qualifications, abilities, and the Company’s needs and resources. The Company does not discriminate in recruiting, hiring, compensation, promotions, discipline, termination or any other aspect of employment on the basis of an individual’s actual or perceived race, color, creed, religion, sex (including pregnancy, childbirth and related medical conditions), sexual orientation, gender identity, national origin, ancestry, citizenship status, age, disability, marital status, military service or status, genetic information, arrest and conviction record, credit history, or any other basis protected by applicable law.
Staff Site Reliability Engineer
360 Privacy
Sr Engineer - Site Reliability
Wwt
Lead Site Reliability & Security Engineer
Clera
Platform Reliability Engineer
Bright Vision Technologies
Senior Site Reliability Engineer (SRE)
LeoLabs, Inc.
Reliability Engineer (Hardware)
Lightmatter