This role provides subject matter expertise and coordination to site reliability efforts across the subdivision. This role also includes ensuring system reliability by meeting service-level objectives (SLOs), driving automation of operational tasks, defining and tracking key performance indicators (KPIs), designing scalable systems, managing incident responses, and collaborating with development teams to ensure software reliability and scalability.
Responsibilites
- Coordinates cross-product chaos experimentation.
- Maintains the centralized incident response playbook for the subdivision to document standards for managing communication and escalation during an incident. Aggregates quantifiable data about availability to report back to senior leadership. Makes contributions to centrally managed (IT-wide) inner source libraries for reliability.
- Facilitates blameless post-incident reviews for high severity incidents or incidents involving more than one product family.
- Regularly attends Reliability Engineering and Resilience communities of practice.
Remains informed about site reliability engineering activities happening within the subdivision.
- Communicates recent standards and newly available tools and frameworks across subdivisions. Enforces reliability standards.
- Participates in special projects and performs other duties as assigned.
- Drives automation of routine operational tasks to improve system efficiency, reduce manual intervention, and enhance deployment and monitoring workflows.
- Leads incident response efforts by diagnosing root causes rapidly, applying timely fixes, and establishing preventive measures to avoid recurrence.
- Defines and tracks key system performance metrics such as availability, latency, and error rates to evaluate and optimize system health and reliability.
- Collaborates with development teams to align software architecture with reliability and scalability goals, ensuring seamless operations across the deployment life