22 Sep
|
Experian
|
Hyderabad
22 Sep
Experian
Hyderabad
Job Description
- Leadership & Strategy
- Define and implement SRE best practices across the organization.
- Proven expertise in production support, engineering, disaster recovery (DCR), automation, and cloud operations
- Mentor and guide a team of SREs, fostering growth
- Collaborate with senior stakeholders to align reliability goals with business objectives.
- Reliability & Performance
- Establish SLIs, SLOs, and SLAs for critical services and ensure adherence.
- Drive initiatives to improve system and reduce operational toil.
- Excellent in designing systems that detect and remediate issues without manual intervention – Self Healing systems, Runbook automation
- Exposure to tools like Gremlin, Chaos Monkey, AWS FIS to simulate outages and improve fault tolerance
- Incident Management
- Act as the primary point of escalation for critical production issues and lead major incident response, root cause analysis, and postmortems.
- Perform detailed post-incident investigations to identify underlying causes.
Document findings and share learnings to prevent recurrence.
- Implement preventive measures and continuous improvement processes.
- Observability
- Champion monitoring, logging, and alerting strategies using tools like Prometheus, Grafana, ELK, and AWS CloudWatch.
- Build real-time dashboards to visualize system health and reliability metrics.
- Configure intelligent alerting based on anomaly detection and thresholds.
- Combine metrics, logs, and traces to enable root cause analysis and reduce Mean Time to Resolution (MTTR).
- Knowledge of AIOps or ML-based anomaly detection for proactive reliability management.
- Collaboration
- Work closely with development teams to integrate reliability into application design and deployment
- Promote a culture of shared responsibility for uptime and performance across engineering teams.
📌 Site Reliability Engineering Lead (Hyderabad)
🏢 Experian
📍 Hyderabad