Senior Technical Lead (India)

Senior Technical Lead (India)

15 Sep
|
HCL Technologies
|
India

15 Sep

HCL Technologies

India

Job Summary

Looking for an extensively experienced Senior Observability Engineer (7-10 years) to lead the design, implementation, and scaling of our enterprise monitoring, telemetry, and observability platform. The successful candidate will be responsible for providing deep system visibility across our infrastructure and microservices running on Amazon EKS (Kubernetes) and AWS. You will join an elite team of platform engineers to formulate observability strategies, eliminate alert fatigue, and ensure maximum reliability across non-prod and production environments.

Key Responsibilities

- Design, build, and maintain scalable observability pipelines (Metrics, Logs, Traces) across AWS and Amazon EKS clusters.
- Establish enterprise-wide standards, guidelines, and governance for telemetry instrumentation using OpenTelemetry (OTel), Prometheus, Grafana, and Zabbix.
- Develop and scale High-Availability monitoring infrastructure, including active-active Zabbix proxy setups and distributed time-series databases.
- Partner with DevOps and Development teams to instrument microservices running in Kubernetes using Helm, Terraform, and automated CI/CD pipelines.
- Integrate observability guardrails into the existing CI/CD pipeline to automate alerting deployment, dashboard provisioning, and health checks.
- Drive SLO/SLI frameworks and Error Budget tracking across product teams to align engineering efforts with business reliability targets.
- Design symptom-based alerting strategies to drastically reduce alert fatigue and improve Mean Time to Resolution (MTTR).
- Troubleshoot complex cross-cutting system issues, latency bottlenecks, and telemetry pipeline failures across hybrid cloud infrastructure.
- Collaborate with Security and Compliance teams to enforce PII/SPI data scrubbing within log aggregation pipelines (e.g., FluentBit, Vector, or Logstash).
- Write and maintain clear architectural documentation, runbooks, and telemetry onboarding guides for product engineering teams.

Skill Requirements

- Bachelor's or Master's degree in Computer Science,



Information Systems, Engineering, or a related field.
- 5-10 years of hands-on experience in Systems, SRE, Platform, or Observability Engineering.
- Strong hands-on expertise in Amazon EKS (Kubernetes), container monitoring, and microservice telemetry.
- Deep experience with 2 or more core observability tools in: Prometheus, Grafana, Zabbix (HA setup), Splunk, New Relic, OpenTelemetry, and APM solutions.
- Proficiency in deploying observability agents and collectors into EKS clusters using Helm and DaemonSets.
- Extensive experience with AWS Core Services: EKS, EC2, VPC, S3, CloudWatch, CloudTrail, Kinesis, and Lambda.
- Strong scripting and automation experience in at least one language: Python, Go, Bash, or Node.js.
- Hands-on experience with Infrastructure as Code (IaC) using Terraform to provision monitoring stacks and dashboards as code.
- Solid understanding of log aggregation architectures (FluentBit, Vector, OpenSearch, or Splunk).
- Experience integrating monitoring and alert routing with ITSM/incident response tools (Jira, PagerDuty, ServiceNow, Teams/Slack).
- Strong understanding of network essentials, system administration, and Linux kernel fundamentals.
- Capable of multitasking, prioritizing in a quick-paced environment, and taking end-to-end ownership of platform reliability.
- A highly effective collaborator with excellent technical documentation skills.

Other Requirements

- AWS Certified Solutions Architect (Associate or Professional) or AWS Certified DevOps Engineer.
- Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD).
- Experience with high-cardinality metrics stores (e.g., Thanos, Cortex, Mimir, or TimescaleDB).
- Familiarity with distributed tracing frameworks (Jaeger, Zipkin, or AWS X-Ray) and W3C Trace Context.
- Prior experience in CI/CD pipeline tools (Jenkins, GitLab CI, or Bitbucket Pipelines) to support GitOps observability (ArgoCD/Flux).
- Knowledge of security posture monitoring and compliance tools (Sysdig, Prisma Cloud, Cloud Custodian).
- Experience working in a Scaled Agile (SAFe/Agile) workflow environment.

📌 Senior Technical Lead (India)
🏢 HCL Technologies
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior technical lead (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: senior technical lead (india) / india