18 Aug
|
Accenture
|
Pune
Program Manager Site Reliability Engineering (Cloud Native Platform Team)Role SummaryThe Program Manager will drive day-to-day operations of the Site Reliability Engineering (SRE) team, ensuring alignment with organizational goals for reliability, scalability, and operational excellence. This role requires a strong technical background in SRE practices and proven program management expertise to drive cross-functional initiatives, optimize processes, and deliver measurable business and operations value.
Key Responsibilities1.Operational LeadershipoDrive adoption of SRE best practices such as error budgets, SLIs/SLOs, and automation to reduce toil.oEnsure compliance with security, privacy, and regulatory standards in all reliability initiatives.2.Program ManagementoDefine program scope, objectives, and success criteria for reliability initiatives.oDevelop and maintain quarterly roadmaps for SRE projects in collaboration with platform engineering teams.oTrack progress, risks, and dependencies across multiple projects using tools like JIRA and Confluence.oFacilitate communication between SRE, development, and leadership teams to ensure transparency and alignment.3.Performance MeasurementoEstablish and monitor KPIs for reliability and operational efficiency.oPrepare executive dashboards and reports to translate technical metrics into business impact narratives.oLead continuous improvement initiatives based on data-driven insights.4.Stakeholder EngagementoAct as the primary liaison between SRE and other teams (Product,
Engineering and Delivery-SOC).oInfluence decision-making at all levels through clear communication and structured reporting.Performance Measurement ParametersIncident Metrics:oMean Time to Detect (MTTD)oMean Time to Respond (MTTR)oMean Time to Recovery (MTTR)oIncident Frequency and SeverityChange Management:oChange Failure RateoChange Success RateReliability Metrics:oSystem Uptime / AvailabilityoService Level Objective Achievement PercentageOperational Efficiency:oAutomation RateoOn-call Burden ReductionMeasurement Matrix for Leadership PresentationUse a dashboard approach combining:Latency, Traffic, Errors, Saturation.Monthly/Quarterly trends on SLO Incident Heatmaps:Highlighting root causes and resolution times.Business Impact Metrics:Cost savings, risk reduction, and ROI from reliability improvements.Tools:DatadogExperience RequirementsTechnical Background:oPrior hands-on experience as a Site Reliability Engineer or in DevOps roles.oStrong understanding of cloud-native architectures (Kubernetes, microservices, distributed systems).Program Management Expertise:o5+ years in program or technical project management.oProven ability to manage cross-functional initiatives in fast-paced environments.oFamiliarity with Agile methodologies and tools (JIRA, Confluence).Leadership Communication:oExperience presenting technical and operational metrics to executive leadership.oStrong stakeholder management and negotiation skills.Certifications:PMP, Protected, or SRE Foundation SAFe preferred
Qualification
📌 SRE Program Manager (Pune)
🏢 Accenture
📍 Pune