Technical Skills Experience with Hadoop ecosystem: Oozie Workflow Manager, Apache Hive, Apache Spark. Working knowledge of Apache Airflow, Managed File Transfer (MFT), and Jenkins CI/CD pipelines. Familiarity with AWS cloud services (S3, EC2, EMR, CloudWatch) in a data platform context. Experience with ITSM platforms, preferably ServiceNow – ticket management, workflow configuration, SLA tracking. Understanding of data pipeline architectures, batch processing, and CDP (Customer Data Platform) environments. Exposure to Agentic AI / GenAI-assisted operations tools is a strong advantage. Operational & Leadership Skills 5+ years of experience in IT/Data managed services or NOC environments, with at least 2 years in a team lead or coordinator role. Demonstrated ability to manage rotational 24×7 support teams across multiple time zones. Strong analytical skills for trend analysis, root cause investigation, and SLA reporting. Excellent written and verbal communication; confident presenting to client stakeholders. Proficient with monitoring and alerting platforms (Grafana, Splunk, CloudWatch, or equivalent). Preferred Qualifications ITIL Foundation certification (v3 or v4). Experience in financial services or fintech data operations. Familiarity with FinOps frameworks and cloud cost monitoring tools. Shift & Team Governance Manage the rotational shift roster (Shift 1: 7AM–4PM | Shift 2: 3PM–12AM | Shift 3: 11PM–8AM IST) ensuring zero coverage gaps 24×7. Conduct daily shift handover briefings; review and sign off on Daily Handoff Notes covering open issues and critical context. Act as first point of escalation for the L1 analyst team; provide on-call support for critical P1 incidents outside general shift hours. Onboard, coach, and performance-manage L1 Support Analysts; identify training needs and coordinate with the EXL capability team. Incident & SLA Management Own the end-to-end incident lifecycle from ticket intake through escalation and closure; ensure P1 escalation to L2 within 10 minutes. Define, review, and update SLAs/KPIs covering availability, response time, acknowledgement (<15 minutes), and resolution time (99% SLA target). Monitor real-time SLA compliance dashboards; intervene proactively when breach risk is detected. Collaborate with Client’s L2 and L3 engineering teams to coordinate resolution of complex incidents and facilitate Root Cause Analysis (RCA). Reporting & Stakeholder Communication Produce Weekly Operational Summaries covering trend analysis, recurring issues, volume metrics, and improvement recommendations. Produce Monthly Performance Reviews tracking SLA compliance, incident metrics (MTTA, MTTR, FCR), and team effectiveness. Participate in Weekly and Quarterly governance calls with Client stakeholders; present insights and action plans. Manage vendor communication and change request tracking; coordinate release planning and operational acceptance sign-off. Knowledge & Continuous Improvement Oversee runbook quality; ensure L1 runbooks and Operational Support Documents (OSDs) are maintained and accessible. Identify automation, tuning, and rationalization opportunities; work with the EXLdata.ai team to implement Agentic AI improvements. Monitor FinOps metrics; provide advisory input on cloud cost optimization and sustainability. Drive First Contact Resolution (FCR) rates from the baseline 45% toward steady-state targets through knowledge base enrichment and pattern-based resolution. Graduate in Computer Science, B.Tech, B.E. 6+years of hands-on data engineering experience.\ 5+ years of experience in IT/Data managed services or NOC environments, with at least 2 years in a team lead or coordinator role. ITIL Foundation certification (v3 or v4). Experience in financial services or fintech data operations. Familiarity with FinOps frameworks and cloud cost monitoring tools.