Gen Digital Inc. logo

Senior Site Reliability Engineer

Gen Digital Inc.
Posted 1 hour ago
MalaysiaHybridEngineering & Development
Is this job info correct?

About Gen:

Gen is a global company dedicated to powering Digital Freedom through its trusted consumer brands including Norton, Avast, LifeLock, MoneyLion and more. Our combined heritage is rooted in financial empowerment and cyber safety for the first digital generations, and today we deliver award-winning cybersecurity, online privacy, identity protection and financial wellness solutions to nearly 500 million users in more than 150 countries.

Together, we share a collective passion and vision to protect consumers and help them grow, manage and secure their digital and financial lives. We’re always looking for smart, fearless and high-impact talent who see AI as a teammate – leveraging it to move faster and deliver meaningful results.

When you’re part of Gen, you’ll have the flexibility, tools and support to do your best work and grow your career – from flexible working options and time off to competitive pay, benefits and well-being programs.

At Gen, we are scrappy and relentlessly customer driven. We create room for healthy debate, experimentation and continuous learning, and we seek out people with different experiences, identities and ideas to join our team. You’ll work with people who back each other, respect each other and understand that our differences are a competitive advantage.

If this sounds like you, we’d love you to be part of Gen.

About the Role

The Kuala Lumpur office is the technology powerhouse of MoneyLion. We pride ourselves on innovative initiatives and thrive in a fast paced and challenging environment. Join our multicultural team of visionaries and industry rebels in disrupting the traditional finance industry!

As a Senior Site Reliability Engineer, you will have the opportunity to build, operate, and evolve the infrastructure and platforms that power MoneyLion’s business-critical applications. You will work across cloud infrastructure, Kubernetes, networking, CI/CD, observability, reliability, and automation to ensure our platforms remain highly available, scalable, secure, and efficient as the business continues to grow.

You will partner closely with Software Engineers, Platform Engineers, Security teams, and other technology teams across MoneyLion to improve reliability and developer experience. As a senior member of the SRE team, you will also lead infrastructure initiatives, drive automation and platform improvements, mentor engineers, and help establish engineering and operational best practices across the organization.

Key Responsibilities

  • Design, build, operate, and continuously improve highly available and scalable infrastructure supporting MoneyLion’s business-critical applications.

  • Build and operate cloud infrastructure and Kubernetes platforms, with a strong focus on reliability, scalability, security, and operational efficiency.

  • Develop and improve Infrastructure as Code, CI/CD platforms, GitOps workflows, and developer self-service capabilities to enable engineering teams to deliver software safely and efficiently.

  • Partner with engineering teams to design reliable and scalable architectures for new and existing applications and services.

  • Improve observability across applications and infrastructure through effective monitoring, alerting, logging, tracing, and automation.

  • Participate in the SRE on-call rotation, lead response to critical production incidents, perform root cause analysis, and drive corrective and preventive improvements.

  • Identify opportunities to reduce operational toil through automation, AI-assisted operations, and improvements to engineering workflows.

  • Lead infrastructure modernization, platform migration, and technology standardization initiatives across MoneyLion.

  • Improve the resilience of MoneyLion’s platforms through capacity planning, performance engineering, disaster recovery, business continuity, and reliability testing.

  • Manage and optimize cloud infrastructure usage and costs while maintaining appropriate levels of performance, reliability, and scalability.

  • Collaborate closely with Security and Engineering teams to ensure infrastructure and platforms follow security, compliance, and operational best practices.

  • Mentor junior engineers, conduct technical and code reviews, and help establish engineering standards and best practices across the SRE and wider engineering organization.

About You

  • Strong software and systems engineering fundamentals, with experience building and operating large-scale production systems in an enterprise environment.

  • Strong hands-on experience with AWS and cloud-native infrastructure, including designing and operating highly available production environments.

  • Demonstrated practical experience with Kubernetes, preferably Amazon EKS, with a strong understanding of Kubernetes architecture, networking, workloads, security, and operations.

  • Strong experience with Infrastructure as Code, preferably Terraform, and experience managing infrastructure through version-controlled and automated workflows.

  • Experience with CI/CD and GitOps practices using platforms such as GitHub Actions or similar technologies.

  • Hands-on experience with observability and monitoring platforms such as Datadog, Prometheus, Grafana, or similar tools.

  • Good understanding of networking concepts including DNS, load balancing, proxies, firewalls, ingress, and cloud networking.

  • Proficiency in scripting or programming languages such as Python, Go, or Bash, with the ability to build automation and operational tooling.

  • Experience troubleshooting complex distributed systems and resolving production incidents across applications, infrastructure, and networking layers.

  • Strong understanding of Site Reliability Engineering practices, including SLOs, SLIs, incident management, capacity planning, disaster recovery, and reducing operational toil.

  • Experience evaluating and adopting new technologies, automation, and AI-assisted engineering tools to improve platform reliability and engineering productivity.

  • Strong problem-solving, communication, and collaboration skills, with the ability to work effectively across multiple engineering and business teams.

  • Experience leading infrastructure or platform initiatives and mentoring engineers, with the ability to take ownership of complex technical problems from design through production operation.

What’s Next

Our interview process is designed to assess both your technical expertise and how you approach real-world reliability and infrastructure challenges.

  • Initial Screening – Virtual (45 minutes)
    An introductory discussion to understand your experience, technical background, and alignment with the Senior SRE role.

  • Technical & Hiring Interview – In Person ( 2.5 hours)
    A face-to-face interview covering technical discussions, a hands-on coding assessment, and a conversation with the Hiring Manager.

__________

Gen is an equal opportunity employer, and we’re committed to fair, inclusive practices at every stage of the candidate and employee journey. Employment decisions are based on merit, experience and business needs.

Similar jobs