Senior System Administrator
- Salary
- CA$90K–CA$95K
- Hiring from
- Canada
- Work type
- Hybrid
- Posted
512,982 remote jobs, straight from company career pages
100% free · New jobs every hour
Show job descriptionHide job description
Data. Discovery. Better Health.
ICES is a world-leading health research and analytics institute. With a wealth of data and analytic expertise, we create trusted evidence that has changed health policy and practice and helps ensure better health for all.
Ready to discover more with us? Join our outstanding, collaborative team where your skills, knowledge and curiosity are valued and can change the future of health care.
At ICES, we recognize what matters most to our employees. Some of the great benefits of working at ICES include:
- Flexible work arrangements
- Competitive Compensation
- Comprehensive Benefits Program
- HOOPP Pension Plan (Defined Pension)
- Employee Assistance Program and Dialogue Well Being Program
- Generous vacation, float and caregiver days for all employees
- Education Fund and Dedicated Education Days
- Holiday Closure
- Perkopolis Employee Discount Program
Introduction:
ICES is seeking a Senior System Administrator to join our Technology department. The Senior System Administrator is responsible for building, maintaining, and supporting Information Technology (IT) infrastructure and systems to ensure optimal performance and reliability. Reporting to the Senior Manager of IT Infrastructure & Operations, this position plays a critical role in managing the organization's technology environment, enabling efficient operations and supporting the overall strategic goals of the organization. By implementing best practices in system administration, the Senior System Administrator contributes to the stability and security of IT services, ultimately enhancing productivity and user satisfaction.
Responsibilities of the position include, but may not be limited to:
Technical Support for IT Infrastructure:
- Provide technical support for High-Performance Computing (HPC) infrastructure and Linux-based systems by utilizing advanced troubleshooting skills and industry knowledge to resolve issues and ensure optimal performance. This includes administering and maintaining HPC cluster environments, including compute, login, and management nodes, while collaborating with IT team members, researchers, analysts, and data scientists to support research and analytics platforms, including Anaconda Python/R environments, Posit Workbench, SAS Grid, and SAS Data Quality Services (DQS) servers;
- Provide technical support to on-prem and Azure-hosted IT infrastructure and systems by utilizing advanced troubleshooting skills and industry knowledge to resolve issues and ensure optimal performance. This includes configuring servers and managing network settings, collaborating with IT team members to assess infrastructure needs and implement solutions that enhance system performance and security;
- Work with documentation and monitoring tools to maintain accurate records of system configurations and changes. This ensures that all technical support aligns with organizational goals and facilitates efficient communication among team members regarding infrastructure status and updates.
Operational Activities Management:
- Perform day-to-day operational activities, including Linux server administration, RHEL patching and upgrades, backup and recovery, infrastructure maintenance, and data center operations by following established protocols and best practices. This includes managing Red Hat Satellite and Red Hat Insights to support system provisioning, lifecycle management, compliance reporting, configuration management, patch management, vulnerability remediation, and system health monitoring;
- Configure and support the Slurm Workload Manager job scheduling environment to ensure efficient workload execution, resource allocation, and system utilization.
- Utilize GitLab repositories for version control, change management, testing, and deployment of infrastructure-as-code and automation workflows, ensuring adherence to established development and operational standards;
- Perform day-to-day operational activities, including server building and administration, system patching and upgrades, backup and recovery, and data center management by following established protocols and best practices. Utilize monitoring software to proactively track system performance and availability, responding swiftly to any issues that arise to minimize downtime;
- Collaborate with other IT staff to coordinate maintenance schedules and ensure that all systems are secure, up-to-date, and functioning efficiently. This involves using project management tools to track operational tasks and ensure timely completion of maintenance activities.
System Monitoring and Troubleshooting:
- Proactively monitor system logs and performance by using analytical skills and monitoring software to identify potential issues before they escalate. This involves troubleshooting and identifying root causes of problems by analyzing system data and logs, developing and implementing technical resolutions for systems hosted on the data center and public Cloud;
- Document findings and solutions to enhance the knowledge base and improve future troubleshooting efforts. This ensures that all team members have access to relevant information and best practices, facilitating continuous improvement in system management.
Collaboration on System Architecture:
- Participate in and provide input to system architecture design, planning, and implementation of new systems by collaborating with cross-functional teams, including developers and project managers. Assess current infrastructure capabilities and recommend enhancements that align with business needs by utilizing knowledge of emerging technologies and industry trends.
Mentorship and Guidance:
- Provide mentorship to the IT infrastructure & operations team by offering guidance and support in technical functions, fostering professional development through regular feedback sessions and access to training resources;
- Assist in departmental projects and initiatives as required by utilizing leadership skills to promote collaboration and alignment with strategic goals, while engaging with cross-functional teams to leverage diverse expertise;
- Foster a culture of continuous improvement by identifying training needs and organizing professional development opportunities for team members, ensuring relevance and effectiveness by utilizing assessment tools and feedback mechanisms.
Other:
- Other duties as may be assigned within the scope of this position.
Knowledge, skills, and abilities:
Minimum Required Education, Years of Experience, and Certifications/Professional Designations:
- Bachelor’s degree in computer sciences or equivalent.
- years progressive system administration experience.
- 5 years Linux administration, and 3 years HPC support experience.
- Experience supporting High-Performance Computing environments within academic, healthcare research, scientific research, or other research-intensive organizations
- IT industry certification required such as Red Hat Certified Engineer (RHCE), Red Hat Certified System Administrator (RHCSA), Red Hat Certified Specialist in Ansible Automation, Linux Foundation Certifications
- Experience with Windows Server and Microsoft Azure administration and deployment
- Microsoft certifications
Minimum Required Position-Specific Knowledge and Technical Skills:
- Strong working experience with RHEL, Ansible Automation Platform, Red Hat Satellite/Insights, GitLab, Slurm, Anaconda Python/R, Posit Workbench, SAS Grid, SAS DQS, Bash/Python scripting, Citrix, NetApp, UCS, VMware.
- Solid understanding and knowledge of Cybersecurity.
- Ability to provide guidance to Help Desk resources, ensuring that support staff are equipped with the necessary knowledge and tools to resolve technical issues effectively and efficiently.
- Strong documentation skills, enabling the creation and maintenance of clear, concise technical procedures, policies, and standard processes related to on-prem and cloud infrastructure.
- Excellent troubleshooting and problem-solving skills, allowing for the identification and resolution of complex technical issues in a timely manner, minimizing downtime and maintaining system reliability.
Preferred Position-Specific Knowledge and Technical Skills:
- Ability to effectively prioritize and execute tasks/projects in a high-pressure environment, demonstrating resilience and focus while managing multiple responsibilities.
- Strong communication and customer service skills, facilitating effective collaboration with team members and partners and ensuring that technical information is conveyed clearly to non-technical audiences.
This full-time, 12-month contract, opportunity is for an existing vacancy at ICES Central. The annual salary range for this role is $90,000 - $95,000.
Security clearance may be required.
Interested candidates should submit their resume and cover letter detailing how their knowledge, skills and abilities match the scope of this position.
ICES is committed to ensuring equity in employment. Our goal is to attract, develop, and retain highly talented employees from diverse backgrounds allowing us to benefit from a wide variety of experiences and perspectives. ICES strongly encourages applications from candidates from equity-deserving communities including but not limited to First Nations, Métis, Inuit, Black and racialized, 2SLGBTQIA+, and persons with visible and non-visible disabilities.
ICES is committed to providing accessible employment practices, in compliance with the Accessibility for Ontarians with Disabilities Act, 2005 (AODA). Applicants are asked to make accommodation requests to ICES and we will make every effort to ensure that accommodation requests are met throughout the recruitment process.
We thank all applicants for their interest in working at ICES. Due to the volume of applications received, only applicants being considered for the position will be contacted for further discussions.