Company Description Tachyon Technologies is a Digital Transformation consulting firm that partners with organizations to deliver customer-focused business transformation. Aligned with SAP’s digital core, the company helps clients leverage existing IT investments and leading-edge digital solutions to improve customer experience and business performance. Tachyon supports clients from initial strategy through full realization, focusing on practical, impactful outcomes. The firm is committed to providing meaningful solutions that exceed client expectations and positions itself as #DigitalIntegrators for modern enterprises.
Designation: Hadoop Support Engineer – L1/L2
Required Experience : 4+ years
Location: Remote
Mode of Hiring: Permanent.
Workdays: Monday to Friday
Work Timings: 6:30 PM - 4:30 AM IST (night shift)
Position Summary:
We are seeking an experienced Hadoop Support Engineer to join our Big Data Operations team, providing Level 1 (L1) and Level 2 (L2) production support across a large-scale, on-premises Hadoop ecosystem. This role is responsible for the day-to-day health, stability, and performance of a multi-cluster environment spanning approximately 700–800 nodes and ~4 PB of storage, built on Acceldata ODP (Apache Hadoop 3.3.x / 3.2.x). The ideal candidate is a hands-on troubleshooter who can monitor, triage, and resolve incidents across HDFS, YARN, Hive, Spark, Trino, ZooKeeper, Ranger, and Ambari — while collaborating with Unix/infrastructure, security, and application teams to keep production workloads running smoothly.
Key Responsibilities:
Production Support & Incident Management (L1/L2):
· Provide first- and second-line production support for a multi-cluster Hadoop environment (700–800 nodes) running Acceldata ODP Hadoop 3.3.x / 3.2.x.
· Monitor cluster health and proactively identify, triage, and resolve incidents affecting NameNodes, ResourceManagers, DataNodes, NodeManagers, ZooKeeper ensembles, Client/Gateway nodes, and Ambari-managed services.
· Perform root cause analysis (RCA) for recurring incidents and service outages, and drive permanent fixes in partnership with L3/engineering teams.
· Manage and resolve an average monthly volume of L2/L3 incidents, service outages, and routine service requests, meeting defined SLAs.
· Own ticket lifecycle management (logging, triage, escalation, resolution, and closure) using the enterprise ITSM/ticketing platform.
· Participate in an on-call rotation and provide off-hours support for critical production incidents and planned maintenance windows.
Cluster Administration & Platform Operations
· Administer and support HDFS and OneFS-based storage deployments (~80% of clusters use OneFS as the primary HDFS filesystem) across on-premises bare-metal infrastructure, with Dev/QA clusters running on VMs.
· Maintain and support core Hadoop components: HDFS, YARN/ResourceManager, ZooKeeper, and Ambari-managed services, across dedicated per-cluster blueprints with isolated resource allocation.
· Support master, worker, and edge node configurations (32–160 vCPUs, 250–2000 GB RAM depending on cluster) and coordinate with infrastructure/Unix teams on hardware and OS-level issues.
· Manage capacity, performance tuning, and resource allocation across NameNodes, ResourceManagers, and NodeManagers to prevent contention and ensure workload SLAs.
· Support cluster upgrades, patching, configuration changes, and decommissioning/commissioning of nodes with minimal production disruption.
Compute & Query Engine Support
· Support and troubleshoot Apache Spark job failures, performance issues, and resource allocation problems.
· Support Apache Hive (including LLAP and Tez execution engines), including query performance issues, metastore health, and job failures.
· Provide support for Trino query engine deployments (currently used in select clusters).
· Work with application and data engineering teams to troubleshoot job submission, scheduling, and resource contention issues across YARN queues.
Security & Governance
· Support Apache Ranger for policy management, RBAC, and access governance across the Hadoop ecosystem.
· Manage user and group access via Active Directory / LDAP synchronization; support onboarding, offboarding, and access remediation requests.
· Support Kerberos authentication where enabled and assist in extending secure authentication coverage across additional clusters as required.
· Enforce data access and governance policies in line with Acceldata ODP Apache governance architecture and enterprise security standards.
· Assist with security audits, access reviews, and compliance reporting related to platform governance.
Monitoring, Alerting & Observability
· Monitor cluster and service health using Acceldata Pulse, SolarWinds, and custom cron-based monitoring scripts.
· Respond to alerts and thresholds breaches, escalating to L3/engineering or infrastructure teams as needed.
· Support and improve log aggregation and troubleshooting practices across clusters, given the currently decentralized, cluster-specific logging implementations.
· Contribute to the standardization of monitoring, alerting, and log aggregation practices across the environment.
· Isilon / OneFS Basics: Health monitoring, storage capacity tracking, basic node/drive health ops — all three L1s cross-trained on Isilon
Collaboration & Documentation
· Partner with Unix/Infrastructure, Network, Security, and Application Development teams to resolve cross-functional issues.
· Maintain and update runbooks, SOPs, and knowledge base articles for recurring issues and standard operating procedures.
· Participate in change management processes for production changes, patches, and maintenance activities.
· Provide clear, timely communication and status updates during incidents and outages to stakeholders and leadership.
Required Qualifications
· Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field (or equivalent practical experience).
· 3–6+ years of hands-on experience supporting or administering Hadoop environments in a production capacity.
· Solid working knowledge of Hadoop 3.x architecture, including HDFS, YARN, Resource Manager/Node Manager, and Zookeeper.
· Experience with Apache Ambari for cluster management, monitoring, and service configuration.
· Hands-on experience supporting Apache Spark and Apache Hive (LLAP/Tez) in production.
· Experience with Apache Ranger for RBAC and policy-based access governance.
· Familiarity with Active Directory / LDAP integration for identity and access management in Hadoop ecosystems.
· Strong Linux/Unix systems fundamentals (networking, storage, performance troubleshooting) in an on-premises, bare-metal environment.
· Experience working within an ITSM/ticketing framework, meeting SLA-driven L1/L2 support commitments.
· Strong analytical and root-cause-analysis skills, with the ability to triage and resolve incidents under time pressure.
· Willingness to work rotational shifts and participate in an on-call rotation, including weekends as needed.
Preferred Qualifications
· Experience with Acceldata ODP or similar enterprise Hadoop distributions.
· Experience with OneFS or other scale-out NAS storage integrated as an HDFS-compatible filesystem.
· Exposure to Trino (or Presto) query engine administration and support.
· Experience with Kerberos authentication setup and troubleshooting in distributed systems.
· Experience with monitoring tools such as Acceldata Pulse, SolarWinds, or equivalent enterprise monitoring platforms.
· Scripting experience (Shell, Python, or similar) for automation of monitoring, health checks, and routine operational tasks.
· Experience supporting large-scale, multi-cluster environments (500+ nodes) with high-availability configurations.
· Relevant certifications (e.g., Cloudera/Hortonworks Hadoop Administrator, Linux, or equivalent) are a plus.
Databricks Competency: Familiarity with Databricks workspace operations — job monitoring, cluster health checks, alerting integrations, and basic triage of Databricks Spark job failures before escalating to the SME.
Interested ?
Please drop your updated resume to [email protected]
Engineer II Technical Support Sales Order Management
Emerson
Technical Support Engineer - Fully Remote | Night Shift
Robert Walters
Sr. Systems Engineer (Internal Support)
Atlas Technica
Associate Technical Support Engineer
Spscommerce
CDN Support Engineer I
bunny.net
Technical Support Engineer II
Delinea