13 Aug
|
HCLTech
|
Bengaluru
Bengaluru, Karnataka
Job Summary
In Observability tribe, we have multiple features teams with team members dedicated to production, there are responsible for the reliability, security, and operational excellence of the group Observability ecosystem that powers observability across the organization. The team operates and scales large, production grade Observability platforms, ensuring they are secure, patched, resilient, and highly available. We work closely with security, infrastructure, and development teams to embed SRE principles, proactive vulnerability management, and standardized patching practices across cloud native and hybrid environments.
Key Responsibilities
Site Reliability Engineering & Platform Operations • Own the reliability, availability, and security posture of Observability platforms and their underlying infrastructure. • Administer, harden, and continuously improve Linux-based systems hosting observability/monitoring tooling. • Design and operate patch management strategies (OS, middleware, runtime, and tooling)
with minimal service disruption. • Implement proactive vulnerability detection, prioritization, remediation, and reporting in collaboration with security teams. • Build and maintain monitoring, alerting, and operational runbooks for platform stability and incident response. ________________________________________ • Lead vulnerability management across CI/CD platforms, containers, and supporting infrastructure. • Analyze vulnerability scans (OS, container, dependency) and drive timely remediation aligned with risk and SLAs. • Automate patching workflows using scripting and infrastructure as code wherever possible. • Ensure compliance with security baselines, hardening guidelines, and internal/external audit requirements. • Reduce security toil through automation, standardization, and preventive controls.
Skill Requirements
Administer and support Observability tool tools such as Elastic Cloud E
📌 Track Manager (Bengaluru)
🏢 HCLTech
📍 Bengaluru