The Work We Do Teciem designs, builds, and delivers treasury and capital markets software solutions for financial institutions worldwide. We serve banks of every size and geography, offering the right setup for the right need. Our solutions are designed to replace multiple disconnected systems with one complete, front-to-back platform, helping customers to capture trading and business opportunities quickly, clearly and with control. We cover the entire trading lifecycle, ensuring that everything - from execution to position keeping, to risk management – runs smoothly. With decades of experience and one of the largest, most diverse client bases in the industry, we turn deep industry knowledge into software that covers most asset classes, meets complex real-world treasury and capital market's needs, and adapts as markets evolve. About Kondor UP Kondor UP is the SaaS edition of Kondor , Teciem's flagship treasury management platform. Operated by TCM (TeCIEM) on AWS , Kondor UP delivers treasury capabilities as a fully managed, always-on service - processing real financial value on behalf of tier-1 banks under strict regulatory requirements. The platform runs on Amazon EKS (Kubernetes), backed by Amazon RDS for SQL Server (Multi-AZ), Amazon MSK (Kafka), Amazon MQ , and a comprehensive observability stack (OpenTelemetry, Prometheus, Grafana, Loki, Jaeger). Changes are delivered through GitOps (ArgoCD) and Terraform , with blue/green deployments and feature flags ensuring zero-downtime operations. Segregation of duties, full audit trails, and always-on availability are not edge cases in this domain - they are the baseline. Role Summary As a Senior Site Reliability Engineer (SRE) - SRE Operations , you will be responsible for the reliability, availability, and performance of the Kondor UP production platform. This is an operations-focused engineering role : you own the production system, define and defend SLOs, lead on-call and incident response, and relentlessly drive down toil through automation and engineering. You apply software engineering discipline to operational problems - turning production failures into systemic improvements, building the observability that gives teams real-time insight, and shifting reliability practices left into the delivery lifecycle. The SRE Operations engineer works closely with the InfraOps , AppOps , NetOps , and SecOps teams, and is a senior contributor within the SRE workstream . Key Responsibilities 1. Service Reliability & SLO Management Define, instrument, and own Service Level Indicators (SLIs) , Service Level Objectives (SLOs) , and error budgets across critical Kondor UP services (deal capture, risk engine, market data feeds, API gateway). Maintain the customer-facing SLA commitments; monitor error budget burn rates (2h / 24h windows) and trigger reliability work when burn thresholds are breached. Produce monthly SLA packs and tenant-scoped reliability reports for customer delivery. Drive Production Readiness Reviews (PRR) before new features or services go live. 2. Incident Response & On-Call Participate in the 24/5 on-call rotation as a senior responder; lead P1/P2 incident coordination, remediation, and customer communication. Triage and escalate across the stack (application, platform, network, data) with authority and speed. Facilitate and own blameless post-mortems : structured root-cause analysis, clear action items, tracked follow-through to prevent recurrence. Maintain and continuously improve on-call runbooks, escalation paths, and the incident management playbook. Enforce Segregation of Duties (SoD) and fully auditable change records during and after every production incident. 3. Observability & Alerting Own the end-to-end observability strategy for Kondor UP: metrics (Prometheus + Grafana), distributed tracing (OpenTelemetry / Jaeger), log aggregation (Loki + Fluent Bit), synthetic monitoring , complementing AWS CloudWatch. Design and maintain alerting rules that are actionable and low-noise - minimising alert fatigue while ensuring timely detection of degradation. Build and maintain tenant-aware SLO dashboards (golden signals: latency, traffic, errors, saturation per tenant). Instrument Istio service mesh metrics for east-west traffic observability within EKS clusters. 4. Production Operations & Platform Health Operate and maintain containerised workloads on Amazon EKS - pod health, Helm release management, node group lifecycle (Karpenter), HPA/KEDA tuning. Manage day-2 operations for platform data services: Amazon RDS SQL Server (Multi-AZ, PITR, failover testing), Amazon MSK (Kafka broker health, consumer lag), Amazon MQ , MemoryDB (Valkey) . Enforce Kyverno admission policies and ensure production environments remain compliant with security and resource standards. Debug complex production issues across containers, microservices, service meshes, and data layers - identify root cause, fix, document, and implement preventive measures. 5. Toil Reduction & Automation Identify, measure, and relentlessly eliminate operational toil through automation - scripting (Python, Bash, Go), self-service tooling, and operational runbook automation. Build and maintain internal operational tooling to enable safe, repeatable, auditable production operations. Set and enforce the team’s toil threshold policy : if toil exceeds a defined percentage of engineering time, reliability work takes priority over feature delivery. 6. Infrastructure as Code & GitOps Apply GitOps principles (ArgoCD) for all application and configuration changes in production - no manual changes, every state transition is a git commit. Contribute to and review Terraform modules (IaC) for platform infrastructure - ensuring all changes are version-controlled, peer-reviewed, and auditable. Validate and gate production deployments: blue/green readiness, canary analysis, smoke tests, rollback criteria. 7. Capacity Planning & FinOps Conduct capacity planning - model workload growth per tenant, set resource headroom policy, and validate auto-scaling behaviour (Karpenter, HPA) under load. Run regular load and performance tests ; validate SLO compliance under projected peak traffic. 8. Production Readiness & Resilience Conduct failure-mode analysis and disaster recovery testing (RDS PITR restoration, cross-region failover, AZ failure simulation). Design and run chaos experiments (LitmusChaos or equivalent) to validate resilience assumptions and improve system robustness. Enforce security guardrails in production: secrets management (AWS Secrets Manager / KMS), CVE triage for running workloads, network policy compliance. Collaborate with the SecOps team on incident response procedures involving security events in production. Required Skills & Experience SRE & Production Operations (mandatory) 5+ years in a Site Reliability Engineer or production operations engineering role, operating a SaaS or always-on service at scale. Proven hands-on experience defining and operating against SLIs, SLOs, error budgets, and customer SLAs . Demonstrated incident management experience: 24/5 on-call, structured root-cause analysis, blameless post-mortems. Strong software engineering fundamentals and proficiency in at least one scripting/programming language (Python, Go, or Bash) for automation and operational tooling. AWS & Kubernetes (mandatory) Hands-on experience operating production workloads on AWS : Amazon EKS, EC2, IAM, VPC, S3, RDS, Route 53, CloudWatch. Solid experience with Kubernetes and Helm for orchestration of containerised workloads in production. Working knowledge of Kubernetes networking (CNI, CoreDNS, NetworkPolicy) and Linux/Unix operating systems. Observability Proficiency with Prometheus and Grafana (dashboards, recording rules, alerting). Experience with OpenTelemetry , distributed tracing (Jaeger or Tempo), and log aggregation (Loki or OpenSearch/ELK). Ability to define and instrument meaningful SLIs from application and infrastructure telemetry. IaC & GitOps Experience with Terraform (modules, remote state, AWS provider) for infrastructure changes in production. Familiarity with ArgoCD or equivalent GitOps tooling for continuous delivery and drift detection. Security & Compliance Knowledge of best practices for data encryption (KMS, TLS/mTLS), secrets management, and least-privilege IAM in production. Awareness of audit and compliance requirements for financial services (SoD, change records, DORA, ISO 27001). Nice to Have Experience operating a service mesh (Istio) for traffic management, mTLS, and fine-grained observability. Experience with Kyverno or OPA for admission control and compliance guardrails in Kubernetes. Experience with HashiCorp Vault or AWS Secrets Manager / KMS for secrets management at scale. Familiarity with FinOps tooling and cloud cost optimisation for a multi-tenant SaaS platform. Experience with chaos engineering and disaster recovery practices (LitmusChaos, Gremlin, GameDay exercises). Exposure to treasury or capital markets systems (Kondor, Summit, or equivalent TMS/risk platforms). AWS certifications: Solutions Architect Professional , DevOps Engineer Professional , or SysOps Administrator . Kubernetes certifications: CKA (Certified Kubernetes Administrator) or CKAD . Profile Experience 5+ years SRE or production operations engineering; 3+ years on AWS in a SaaS or always-on context Education Engineering degree (Computer Science, Information Technology, Mathematics, or equivalent) or equivalent experience Languages English (mandatory - working language) Location Bangalore, India Work model Hybrid - Bangalore (TeCIEM India office) + remote On-call Participation in a 24/5 on-call rotation is a core requirement of this role Team Context & Reporting The SRE Operations Engineer (TP3) reports to the Cloud SRE Manager and operates within the SRE workstream . Primary interfaces: - InfraOps - joint ownership of EKS cluster health, node lifecycle, platform baseline; escalation path for infrastructure-layer incidents. - AppOps - first-line production support; SRE provides reliability tooling, runbooks, and post-mortem leadership. - NetOps - network-layer triage during incidents; collaboration on SLI instrumentation for connectivity paths. - SecOps - production security posture, CVE triage, incident response coordination, compliance evidence. - DevOps / Platform Engineering - reliability gates in CI/CD pipelines; PRR process; shared ownership of GitOps toolchain. - Architecture - resilience pattern review, capacity planning inputs, ADR contributions. Diverse Minds, Shared Ambition At Teciem, we believe that our strength comes from the diversity of our people. Different perspectives, backgrounds, and experiences fuel our innovation and help us build solutions that truly make a difference in the world of financial technology. We’re committed to creating a workplace where everyone feels respected, heard, and empowered to grow. Here, you can bring your whole self to work, contribute your unique ideas, and be part of a team driven by shared ambition. We welcome talent from all walks of life and encourage applications from individuals of all genders, races, ages, abilities, identities, and beliefs. Together, we’re shaping a culture where diversity isn’t just celebrated — it’s essential to our success.
Site Reliability Engineer (SRE) Operations
Teciem
Operations Engineer
Syniverse
Infrastructure Operations Engineer
Warnerbros
Site Reliability Engineer/Cloud Platform Engineer - Operations (PST Timezone)
Skyflow
Network Operations & Services Engineer
Joblistings
Senior Technical Recruiter - Software Engineering & Operations
Constantinople