01 Aug
|
Recognized
|
Hyderabad
01 Aug
Recognized
Hyderabad
Lead Observability Strategy: Define and execute the observability roadmap aligned with business and IT goals, integrating AIOps and SRE principles. Tool Ownership & Integration: Manage and optimize observability tools including OpsRamp, Splunk, AppDynamics, NetBrain, ThousandEyes, and explore recent platforms like BigPanda and ServiceNow AIOps. Automation Leadership: Drive automation of L1/L2 operational tasks using Python and PowerShell, improving efficiency and reducing manual intervention. SRE Adoption: Collaborate with cross-functional teams to implement Site Reliability Engineering (SRE) practices, including SLIs/SLOs, error budgets, and incident response automation. Monitoring & Dashboarding: Design and maintain comprehensive dashboards and alerting mechanisms for infrastructure, applications, and network performance. Incident & Problem Management: Partner with ITSM teams to enhance incident detection, root cause analysis,
and resolution workflows. Mentorship & Collaboration: Lead and mentor a team of observability engineers, fostering a culture of innovation, ownership, and continuous improvement. 8+ years of experience in IT operations, observability, or infrastructure monitoring. Strong hands-on experience with tools like Splunk, OpsRamp, AppDynamics, NetBrain, ThousandEyes. Experience with AIOps platforms (BigPanda, ServiceNow AIOps preferred). Proficiency in Python and PowerShell for automation and scripting. Familiarity with SRE principles and implementation strategies. Solid understanding of ITIL processes (Incident, Change, Problem Management). Excellent communication, leadership, and stakeholder management skills.
📌 Observability Lead (Hyderabad)
🏢 Recognized
📍 Hyderabad