Site Reliability Engineer (SRE)
Purpose & Overall Relevance for the Organization:
- Maintain and enhance monitoring framework (data collection, alert aggregation, dashboarding) and implement and enhance alerting logic (framework).
- Enable proactive incident alert and resolution leveraging knowledge scripts.
- Identify and detect repetitive incidents (stability, reliability) and develop solutions to fix problems.
- Work on technical resolution for incidents and identify technical root cause.
- Ensure tool standards, exploit tool capability to fine-tune product reliability.
- Integrate incident, release, monitoring, and alerting tools into the overall ecosystem.
- Measure and report SLI, MTTx in periodic reviews, analyze deviations, and take actions to closure.
- Update runbooks with changes to process/tools.
- Drive postmortems to arrive at remedial actions.
- Participate in On-Call Incident Technical Support.
- Ensure production release guidelines (entry/exit) and implementation are adhered to for changes to Production.
- Support CI/CD pipeline implementation and integration to quality, security.
- Scale systems sustainably through mechanisms like automation; evolve systems by pushing for changes that improve reliability and velocity.
Key Responsibilities:
- Maintain and enhance monitoring framework (data collection, alert aggregation, dashboarding).
- Implement and enhance alerting logic (framework).
- Enable proactive incident alert and resolution leveraging knowledge scripts.
- Identify and detect repetitive incidents (stability, reliability) and develop solutions to fix problems.
- Work on technical resolution for incidents and identify technical root cause.
- Ensure tool standards, exploit tool capability to fine-tune product reliability.
- Integrate incident, release, monitoring, alerting tools into the overall ecosystem.
- Measure and report SLI, MTTx in periodic reviews, analyze deviations, and take actions to closure.
- Update runbooks with changes to pro