Principle SRE (Site Reliability Engineering)
Role Overview:
The Principal Site Reliability Engineer will be a senior technical expert responsible for driving end-to-end resilience, reliability, and scalability across our mission-critical payments platform. This role focuses on front-to-back payment flows, ensuring systems are designed for fault tolerance, observability, and operational excellence.
Key Responsibilities:
- Reliability Engineering Leadership:
- Drive strategies to improve reliability, maintainability, and scalability across payment flows and platform components.
- Architecture & Design Reviews:
- Conduct deep technical assessments of system architectures, identifying risks and recommending improvements for fault tolerance and disaster recovery.
- Incident Management & Root Cause Analysis:
- Act as a senior escalation point for production incidents, lead RCA, and implement permanent fixes to prevent recurrence.
- Resiliency by Design:
- Define and enforce reliability patterns, frameworks, and best practices; ensure adoption across engineering teams.
- Chaos Engineering & Failure Testing:
- Advocate and implement chaos engineering principles to validate system resilience under real-world failure scenarios.
- Observability & Monitoring:
- Design and implement full-stack observability solutions, including metrics, logging, distributed tracing, and alerting.
- Automation & Tooling:
- Develop automation for failover, capacity management, and self-healing mechanisms to reduce operational risk.
- Collaboration:
- Partner with development, infrastructure, and production support teams to embed reliability into the SDLC.
- Continuous Improvement:
- Analyze service risk assessments and production incidents to identify systemic issues and drive long-term improvements.
- Culture Building:
- Promote operational excellence and a mindset of designing for failure across all engineering teams.
Required Skills & Experience:
- Technical Expertise:
- 12+ years in softwar
📌 Principle Sre (Pune)
🏢 Barclays
📍 Pune