Site Reliability Engineer (Cloud Operations)
- Salary
- $100K–$130K
- Hiring from
- United States
- Work type
- Hybrid
- Posted
526,903 remote jobs, straight from company career pages
100% free · New jobs every hour
Show job descriptionHide job description
Concept Solutions, LLC (CS) is seeking a skilled and motivated Site Reliability Engineer (Cloud Operations) to support a mission-critical application modernization program for a Department of War (DoW) customer.
This initiative will unify access across a portfolio of production applications through a single sign-on experience, a common workspace, and a repeatable approach for onboarding applications to a shared platform. The program will also maintain and sustain existing production applications throughout the modernization effort. Contingent Upon Contract Award
In this role, you will support the reliability, availability, and operational performance of 13 containerized services across two AWS accounts. You will work closely with application development teams, platform engineers, and DevSecOps personnel to maintain monitoring and alerting capabilities, coordinate incident response, manage patching activities, and document service availability.
This opportunity is ideal for an engineer who enjoys troubleshooting production systems, improving operational processes, and helping deliver reliable cloud-based services in a mission-focused environment.
Cloud Operations and Monitoring
- Operate and maintain monitoring, alerting, and operational support processes for 13 containerized services across two AWS accounts.
- Monitor application health, service availability, and operational performance to identify and address potential issues.
- Support observability capabilities using monitoring dashboards, logging, metrics, and alerting tools.
- Investigate operational issues, troubleshoot service disruptions, and coordinate resolution with appropriate technical teams.
- Identify opportunities to improve monitoring effectiveness, service stability, and operational efficiency.
Incident Management and Service Reliability
- Maintain and support incident response procedures, escalation processes, and on-call operations for production services.
- Participate in incident investigation, troubleshooting, service recovery, and post-incident reviews.
- Document incidents, identify contributing factors, and recommend improvements to reduce recurring operational issues.
- Maintain operational procedures and documentation to support consistent incident response and service continuity.
- Collaborate with engineering teams to address service reliability risks and improve production support processes.
Patching, Maintenance, and Operational Reporting
- Coordinate and execute patch cycles for container images and application dependencies in partnership with application development teams.
- Support operational maintenance activities to help ensure the continued security, stability, and reliability of production services.
- Track and document patching activities, operational issues, and remediation efforts.
- Produce and maintain service availability evidence used in monthly operational reporting.
- Work with application teams to identify and resolve issues related to containerized application environments.
Team Collaboration and Continuous Improvement
- Collaborate with DevSecOps, platform engineering, and application development teams to support reliable application operations.
- Participate in technical discussions, operational reviews, and continuous improvement activities.
- Contribute to operational documentation, runbooks, and knowledge-sharing efforts.
- Support Agile delivery activities, including two-week sprint cycles and customer demonstrations.
- Minimum of 4 years of professional experience in Site Reliability Engineering (SRE), operations engineering, or a closely related role supporting containerized applications.
- Hands-on experience supporting containerized applications using Kubernetes or Amazon Elastic Container Service (ECS).
- Experience with observability tools and practices, including monitoring, logging, metrics, and alerting.
- Knowledge of incident management processes, production troubleshooting, incident response, and service recovery.
- Experience supporting operational maintenance and troubleshooting of production application environments.
- Ability to collaborate with application development and infrastructure teams to resolve technical issues.
- Strong analytical, problem-solving, documentation, and communication skills.
Preferred Qualifications
- Experience supporting applications and infrastructure within AWS GovCloud.
- Familiarity with OpenTelemetry for application observability, metrics, and distributed tracing.
- Experience supporting cloud-hosted, containerized production services.
- Familiarity with container image patching, dependency maintenance, and operational automation.
- Experience participating in on-call support rotations and post-incident reviews.
- Experience supporting federal government or other security-conscious environments.
- Active Secret security clearance or higher.
- Eligibility for a Tier 5 investigation, if required for privileged administration.
Work Location and Schedule
- Location: Hybrid in the Washington, DC Metro Area
- Schedule: Full-time, 40 hours per week, Monday through Friday, with core hours of 9:00 a.m. to 3:00 p.m. Eastern.
- Security Requirement: U.S. citizenship is required. Candidates must be eligible to obtain and maintain a DoD security clearance. An active Secret clearance or higher is strongly preferred and may allow for an earlier start.
- Compensation: $100K to $130K
Company Profile
Founded in 1999 and headquartered in Reston, Virginia, Concept Solutions, LLC (CS) is a leading small business in technology, engineering, and management consulting. We are the innovative and agile force behind strategic solutions that enhance organizational efficiency and safeguard our nation across Aerospace, Defense, and National Security sectors.
For over 25 years, CS has been a trusted partner for the Federal Aviation Administration (FAA), Department of Homeland Security (DHS), Department of Justice (DOJ), Department of Defense (DoD) and other federal agencies delivering vital IT, security, and project management services.
Our commitment to excellence is reflected in our adherence to CMMI-DEV ML3, ISO 9001:2015, ISO/IEC 20000-1:2018, and ISO/IEC 27001-1:2013 standards. CS boasts company highlights that include:
- Over two decades of experience across over $300 million in contract awards supporting critical FAA programs
- Multiple contract vehicles providing opportunities across FAA, DoD, NOAA, and other Federal agencies
- Innovation Council - CS maintains an active Internal Research and Development (IR&D) program that is geared towards identifying emerging technologies and pursuing technological innovations
At CS, we know our success stems from our talented team. That’s why we prioritize the wellbeing and growth of our employees, fostering a positive culture centered on innovation, engagement, and career development.
Benefits
Concept Solutions offers a competitive benefits and salary package you would receive from a large company. We offer health, dental, vision and life insurance, as well as a comprehensive 401(k) plan with matching and immediate vesting.
Concept Solutions is an Equal Opportunity Employer, and we value workplace diversity. We invite resumes from all interested parties and consider applicants for all positions without regard to race, color, religion, sex, national origin, age, marital status, sexual preference, personal appearance, family responsibility, the presence of a non-job-related medical condition or physical disability, matriculation, political affiliation, veteran status, or any other legally protected status.
Concept Solutions is a VEVRAA federal contractor, and we request priority referral of veterans for available positions.