04 Oct
|
HCLTech
|
Hyderabad
Track Manager (Support &
• Operations) Hyderabad, Telangana
Job Summary
Tools – L3 Engineer (Splunk/DataDog) Key Responsibilities: Architecture &
• Strategy: Architect, design, and implement end-to-end observability solutions across hybrid/multi-cloud platforms (2100 Azure VMs, 139 GCP VMs, Datacenters, GKE/AKS clusters). Advanced L3 Escalations &
• RCA: Serve as the final technical escalation point for critical (Sev-1/Sev-2) production incidents; perform deep-dive root cause analysis (RCA) using distributed tracing, logs, and metrics. AIOps &
• Automation Integration: Drive AIOps-enabled incident, service request (SR), and change (CHG) workflows; automate auto-remediation scripts and golden path monitoring templates. Cost Optimization &
• Lifecycle Management: Architect cost-effective log ingestion pipelines, filtering, sampling, and data retention policies in Datadog and Splunk to optimize license utilization. Infrastructure as Code (IaC): Standardize observability deployments using IaC (Terraform, Ansible) into CI/CD DevSecOps pipelines for zero-touch provisioning of monitoring agents across infrastructure and containers.
Custom
Integration &
• Tooling: Develop custom scripts, APIs, and plugins (Python, Bash) to integrate non-standard databases (Databricks, CosmosDB, DB2) and network assets into unified dashboards. SLO/SLI &
• Governance: Partner with Platform Engineering and Business teams to define, implement, and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets. Mentorship &
• Process Engineering: Mentor L1/L2 operational teams, review standard operating procedures (SOPs), and lead technical engagement for new project onboarding.
Key Responsibilities
Key Responsibilities: Architecture & Strategy: Architect, design, and implement end-to-end observability solutions across hybrid/multi-cloud platforms (2100 Azure VMs, 139 GCP VMs, Datacenters, GKE/AKS clusters). Advanced L3 Escalations & RCA:
Serve as the final technical escalation point for critical (Sev-1/Sev-2) production incidents; perform deep-dive root cause analysis (RCA) using distributed tracing, logs, and metrics. AIOps & Automation Integration: Drive AIOps-enabled incident, service request (SR), and change (CHG) workflows; automate auto-remediation scripts and golden path monitoring templates. Cost Optimization & Lifecycle Management: Architect cost-effective log ingestion pipelines, filtering, sampling, and data retention policies in Datadog and Splunk to optimize license utilization. Infrastructure as Code (IaC): Standardize observability deployments using IaC (Terraform, Ansible) into CI/CD DevSecOps pipelines for zero-touch provisioning of monitoring agents across infrastructure and containers. Custom Integration & Tooling: Develop custom scripts, APIs, and plugins (Python, Bash) to integrate non-standard databases (Databricks, CosmosDB, DB2) and network assets into unified dashboards. SLO/SLI & Governance: Partner with Platform Engineering and Business teams to define, implement, and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets. Mentorship & Process Engineering: Mentor L1/L2 operational teams, review standard operating procedures (SOPs), and lead technical engagement for new project onboarding
Skill Requirements
Technical: Splunk (L3 Level): Expert-level knowledge of Splunk Enterprise architecture, cluster management, complex SPL development, data onboarding, custom technology add-ons (TAs), and enterprise security monitoring. Datadog (L3 Level): Advanced proficiency in Datadog APM, Distributed Tracing, Profiler, Log Management,
Network Performance Monitoring (NPM), and Security Monitoring. Cloud & Containers: Deep expertise in monitoring Multi-Cloud environments (Azure, GCP) and orchestrators (Kubernetes / GKE / AKS). Automation & Scripting: Strong competency in Terraform, Ansible, Python, and Bash for infrastructure automation and API integrations. Observability Concepts: Deep technical understanding of OpenTelemetry (OTel), distributed tracing, AIOps, anomaly detection, and synthetic monitoring. Non-Technical Robust architectural thinking, problem-solving, and crisis management capabilities during complex technical outages. Proven experience in technical leadership, stakeholder management, and cross-team collaboration (Cloud, Network, DBAs, DevOps). Mastery of ITIL framework (Problem Management, Capacity Management, Service Transition). Good to Have Skills: Hands-on experience with secondary enterprise tools like SolarWinds, Apptio, Commvault, or Cisco/SilverPeak SD-WAN monitoring. Exposure to database internals for tuning monitoring parameters (MSSQL, Oracle, PostgreSQL, Databricks). Knowledge of FinOps frameworks for monitoring tool spend optimization. Soft Skills: Executive communication and technical reporting skills. Strong decision-making capability under high-pressure, major-incident situations. Proactive mindset focused on continuous improvement and operational excellence.
Other Requirements
Certifications: Datadog: Datadog Certified Administrator / Datadog Advanced APM & Log Management (Required/Preferred) Splunk: Splunk Enterprise Certified Architect or Splunk Core Certified Consultant (Highly Preferred) Cloud & DevOps: Azure Solutions Architect (AZ-305) / GCP Professional Cloud Architect / Certified Kubernetes Administrator (CKA) (Good to have) Overview Role – Tools – L3 Engineer (Splunk/DataDog) Experience Level – 8 years Job Location – Bengaluru / Hyderabad Job Type – Full-Time Support type - 24x7 Operational Support & Project Engagements
📌 Track Manager (Support & Operations) (Hyderabad)
🏢 HCLTech
📍 Hyderabad