Senior Site Reliability Engineer (SRE)
VtechsolutionsJob Description
This is a remote position.
Location: Remote – United States MUST BE ABLE TO WORK EST HOURS
Our client, a growing technology services organization supporting federal government programs, is seeking an experienced Senior Site Reliability Engineer to support a federal financial agency.
This role is focused on the reliable, secure, and efficient operation of critical production infrastructure. The Senior Site Reliability Engineer will be responsible for maintaining production environments, troubleshooting incidents, supporting deployments, improving monitoring and diagnostics, and driving greater uptime and operational stability.
Responsibilities
- Maintain, monitor, and troubleshoot production environments to support system uptime, reliability, and performance.
- Manage and operate infrastructure using Terraform, Ansible, and Docker.
- Support and maintain CI/CD pipelines and automate operational workflows using Git and related tooling.
- Ensure the ongoing reliability of systems operating within AWS environments, including EKS, S3, and EMR.
- Support operational use of technologies such as Spark, JupyterHub, and Hue.
- Diagnose and resolve infrastructure and application issues.
- Conduct root-cause analysis and drive long-term resolution of recurring incidents.
- Implement, refine, and maintain infrastructure and application monitoring, alerting, and diagnostics.
- Support deployment activities, maintenance windows, and production changes.
- Optimize data flows and storage integrations.
- Collaborate with engineering, product, and client stakeholders to communicate issues, coordinate maintenance, and support operational priorities.
- Contribute to continuous improvement of operational processes, platform documentation, and reliability best practices.
Requirements
- Active Secret security clearance or higher is required
- Strong experience supporting and maintaining production infrastructure.
- Hands-on Python experience within operational, infrastructure, or support environments.
- Professional experience with Terraform, Ansible, and Docker.
- Experience supporting CI/CD pipelines and deployment workflows.
- Strong proficiency with Git and version-control practices.
- Strong Linux command-line and systems operations experience.
- Experience monitoring, diagnosing, and troubleshooting production systems.
- Strong incident-response and root-cause analysis capabilities.
- Ability to anticipate and resolve complex operational issues.
- Strong communication skills and the ability to work effectively in a collaborative, client-facing environment.
- Ability to learn and adapt to new technologies quickly.
- Ability to work East Coast business hours.
Preferred Qualifications
- Experience operating infrastructure within AWS or another major cloud platform.
- Hands-on experience with AWS services including EKS, S3, and EMR.
- Familiarity with Spark, JupyterHub, and Hue in an operational environment.
- Databricks experience.
- Experience supporting federal government, regulated, or other security-sensitive environments.
- Experience working directly with external clients or government stakeholders.
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline.