06 Aug
|
Tachyon Technologies
|
India
06 Aug
Tachyon Technologies
India
Job Title: Lead – L3 Hadoop Support Engineer
Working hours:
30 hours per week. Primary business-hours coverage with participation in a 24x7 on-call escalation rotation for critical (P1/P2) incidents
Level 3 Escalation Lead
— Big Data Platform Engineering & Customer Support
Department:
Enterprise Data Platform / Big Data Operations
Employment Type:
Full-Time
Work Model:
On-Premises Data Center Environment
Position Summary:
We are looking for a Lead Hadoop Support Engineer to serve as the Level 3 (L3) technical authority and team lead for our Big Data Operations group. This role sits above L1/L2 production support and owns the hardest technical problems: deep root-cause investigation, cross-cluster performance and capacity issues, security/architecture decisions, and direct engagement with internal customers and application teams on complex requests. You will mentor and guide the L1/L2 support team, act as the final escalation point before vendor/engineering involvement, and be accountable for the stability of a large-scale, on-premises Hadoop ecosystem spanning approximately 700–800 nodes and ~4 PB of storage, built on Acceldata ODP (Apache Hadoop 3.3.x / 3.2.x).
Strong hands-on expertise across HDFS, YARN, Hive, Spark, Trino, ZooKeeper, Ranger, and Ambari is essential, as is the ability to communicate confidently with both engineers and business stakeholders.
Key Responsibilities:
L3 Escalation Ownership & Customer-Facing Technical Support
Serve as the final internal escalation point for incidents and requests that L1/L2 cannot resolve, including complex, cross-component, or previously unseen production issues.
Own and lead root cause analysis (RCA) for major incidents and service outages, producing clear, business-appropriate RCA reports and permanent corrective actions.
Act as the primary technical point of contact for internal customers and application teams on complex data platform requests — translating business/customer requirements into technical solutions.
Personally handle and resolve customer-reported technical queries that require deep platform expertise (query performance, job failures, data access issues, capacity constraints, security exceptions).
Own communication with stakeholders and leadership during major (P1/P2) incidents, providing accurate technical updates and realistic resolution timelines.
Represent the support function in post-incident reviews, change advisory board (CAB) meetings, and customer/stakeholder review calls.
Team Leadership & Technical Mentorship
Provide day-to-day technical guidance, coaching, and mentorship to L1/L2 support engineers, raising the team's overall troubleshooting capability.
Review and approve L1/L2 incident resolutions and RCAs for technical accuracy before closure on complex tickets.
Define, document, and continuously improve runbooks, troubleshooting guides, and standard operating procedures (SOPs) for the team.
Support hiring, onboarding, and technical skills development (training plans, shadowing, knowledge-transfer sessions) for new and existing team members.
Plan and coordinate on-call escalation rotations,
ensuring adequate L3 coverage and smooth handoffs between L1/L2 and L3.
Conduct regular technical reviews with the team on recurring issues, ticket trends, and opportunities for automation or permanent fixes.
Platform Architecture, Performance & Capacity
Own performance tuning and capacity planning across NameNodes, ResourceManagers, DataNodes, NodeManagers, and ZooKeeper ensembles for a 700–800 node, multi-cluster environment.
Lead architecture and configuration decisions for HDFS and OneFS-based storage deployments, and for master/worker/edge node sizing (32–160 vCPUs, 250–2000 GB RAM per cluster).
Drive cluster upgrade planning, version compatibility assessments, and major configuration changes (Hadoop 3.3.x/3.2.x, Acceldata ODP) with minimal production risk.
Own end-to-end troubleshooting and optimization of Apache Spark, Apache Hive (LLAP/Tez), and Trino workloads, including query tuning and resource-queue design.
Evaluate and recommend improvements to monitoring and observability tooling (Acceldata Pulse, SolarWinds, custom scripts) and drive standardization of log aggregation across clusters.
Partner with Acceldata and other vendor support teams on complex defects, patches, and platform roadmap items.
Security, Governance & Compliance
Own the technical design and enforcement of Apache Ranger policies, RBAC models, and access governance across all clusters.
Lead efforts to standardize and extend Kerberos authentication coverage across clusters where not yet enabled.
Oversee Active Directory / LDAP integration, access provisioning standards, and periodic access/security reviews.
Ensure the platform's data governance model aligns with the Acceldata ODP Apache governance architecture and enterprise compliance requirements.
Support audits by providing technical evidence, remediation plans, and sign-off on security-related findings.
Process, Reporting & Continuous Improvement
Analyze monthly ticket trends (typically 5–10 L2/L3 incidents, 0–1 outages, and 10–20 routine requests) to identify systemic issues and drive permanent fixes that reduce recurring volume.
Own and continuously improve the incident, problem, and change management processes for the Hadoop platform in partnership with ITSM/process owners.
Define and track SLA/SLO adherence for L1/L2/L3 support, reporting on trends to platform leadership.
Identify opportunities for automation of monitoring, health checks, and routine operational tasks to reduce manual L1/L2 effort.
Required Qualifications
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field (or equivalent practical experience).
7+ years of hands-on Hadoop administration/support experience, including 2+ years in a lead, senior,
or L3 escalation capacity.
Deep expertise in Hadoop 3.x architecture — HDFS, YARN, ResourceManager/NodeManager, ZooKeeper — including advanced troubleshooting and performance tuning.
Proven experience owning production Apache Spark and Apache Hive (LLAP/Tez) issues end-to-end, including query and resource optimization.
Strong working knowledge of Apache Ranger (RBAC/policy design) and Active Directory / LDAP integration in enterprise Hadoop environments.
Experience with Apache Ambari for cluster administration, and familiarity with Kerberos authentication design and troubleshooting.
Demonstrated experience acting as a technical point of contact for internal or external customers, including explaining technical issues to non-technical stakeholders.
Experience mentoring, coaching, or formally leading a technical support team (L1/L2), including reviewing others' work and building team capability.
Strong Linux/Unix systems expertise (networking, storage, performance) in large-scale, on-premises, bare-metal environments.
Excellent root-cause-analysis, incident command, and written/verbal communication skills, with composure under P1/major-incident pressure.
Willingness to participate in an on-call escalation rotation, including occasional off-hours involvement for critical incidents.
Preferred Qualifications
Experience with Acceldata ODP specifically, or comparable enterprise Hadoop distributions (Cloudera, Hortonworks).
Experience with OneFS or other scale-out NAS storage integrated as an HDFS-compatible filesystem, at multi-petabyte scale.
Hands-on experience administering or tuning Trino (or Presto) in production.
Experience with Acceldata Pulse, SolarWinds, or similar enterprise monitoring/observability platforms.
Scripting/automation experience (Shell, Python) for operational tooling, health checks, and self-healing automation.
Prior experience running or contributing to Change Advisory Board (CAB) processes and formal incident/problem management.
Experience supporting environments of 500+ nodes with multi-cluster isolation and high-availability design.
Relevant certifications (e.g., Cloudera/Hortonworks Hadoop Administrator, ITIL Foundation, Linux, or equivalent).
Databricks Competency: Familiarity with Databricks workspace operations — job monitoring, cluster health checks, alerting integrations, and basic triage of Databricks Spark job failures before escalating to the SME
What Success Looks Like in This Role
L1/L2 escalations are resolved faster and more consistently because the team has explicit guidance, better runbooks, and a strong L3 backstop.
Recurring incident volume trends down month over month due to permanent fixes rather than repeated workarounds.
Internal customers and application teams see the Hadoop support function as a trusted, responsive technical partner — not just a ticket queue.
Security, governance, and monitoring practices become more standardized and consistent across all clusters.
Duties, responsibilities, and requirements may change at any time with or without notice.
📌 Lead Hadoop Engineer(9+ years) (India)
🏢 Tachyon Technologies
📍 India