10 Aug
|
HCLTech
|
Bengaluru
Key Responsibilities
- Implement and institutionalize SRE practices including SLIs, SLOs, Error Budgets, reliability reviews, and service health management
- Design and implement end-to-end observability solutions using Splunk Observability Cloud, Splunk ITSI, and Splunk Enterprise
- Define and manage alert quality standards to reduce noise, eliminate alert fatigue, and improve incident response effectiveness
- Configure and maintain business service monitoring, KPI frameworks, service health dashboards, event analytics, and business observability capabilities
- Implement AIOps capabilities including event correlation, anomaly detection, intelligent alerting, and operational analytics
- Drive continuous improvement initiatives to improve MTTD, MTTR, MTBF, service reliability, and operational efficiency
- Integrate monitoring, logging, ITSM, cloud, application, infrastructure, and third-party platforms to enable end-to-end observability
- Support incident management, problem management, RCA, and blameless postmortem activities
- Partner with application, cloud, platform, and operations teams to improve production readiness and operational resilience
- Develop automation solutions for monitoring, alerting, reporting, remediation, and operational workflows
- Participate in Agile ceremonies, reliability reviews, and continuous service improvement programs
Skill Requirements
Must Have Skills
- 8+ years of experience in Site Reliability Engineering, Production Support, Application Support, Observability, Operations Engineering, or related roles
- Strong hands-on experience implementing SRE practices including SLI, SLO, Error Budgets, reliability governance, and service health management
- Strong expertise with Splunk Observability Cloud (OpenTelemetry, Infrastructure Monitoring, APM, RUM, Synthetic Monitoring)
- Strong expertise with Splunk ITSI (Event Analytics, Service Modeling, KPI Management, Glass Tables, Service Health Monitoring, Business Observability)
- Strong expertise with Splunk Enterprise (Logging, SPL, Log Analytics, Dashboarding)
- Experience implementing alert quality management, alert rationalization, and noise reduction initiatives
- Experience with Incident Management, Problem Management, RCA, and Blameless Postmortems
- Experience implementing AIOps and intelligent operations capabilities
- Hands-on experience integrating observability platforms with applications, cloud services, infrastructure, ITSM platforms, and enterprise tools
- Experience supporting business applications including custom applications (Java, .NET and other technologies running on VM and container platforms), SAP, Salesforce, and SaaS/COTS platforms
- Experience with automation technologies to reduce the operational toil (Ansible, Python, RPA and other automation platforms)
- Robust analytical, problem-solving, communication, and stakeholder management skills
Other Requirements
Good to Have Skills
- Knowledge of DevOps practices, CI/CD pipelines, GitOps, and release automation
- Experience with Platform Engineering and Internal Developer Platforms (IDP)
- Experience with Resilience Engineering and Chaos Engineering practices
- Exposure to Agentic AI, AI-driven Operations, and AI-assisted observability solutions
- Experience with cloud platforms including Azure, AWS, and GCP
- Splunk, SRE, Cloud, or Observability-related certifications
📌 SRE Expert (Bengaluru)
🏢 HCLTech
📍 Bengaluru