Principle SRE (Site Reliability Engineering)
Role Overview:
The Principal Site Reliability Engineer will be a senior technical expert responsible for driving end-to-end resilience, reliability, and scalability across our mission-critical payments platform. This role focuses on front-to-back payment flows, ensuring systems are designed for fault tolerance, observability, and operational excellence.
Key Responsibilities:
Reliability Engineering Leadership:
Drive strategies to improve reliability, maintainability, and scalability across payment flows and platform components.
Architecture & Design Reviews:
Conduct deep technical assessments of system architectures, identifying risks and recommending improvements for fault tolerance and disaster recovery.
Incident Management & Root Cause Analysis:
Act as a senior escalation point for production incidents, lead RCA, and implement permanent fixes to prevent recurrence.
Resiliency by Design:
Define and enforce reliability patterns, frameworks, and best practices; ensure adoption across engineering teams.
Chaos Engineering & Failure Testing:
Advocate and implement chaos engineering principles to validate system resilience under real-world failure scenarios.
Observability & Monitoring:
Design and implement full-stack observability solutions, including metrics, logging, distributed tracing, and alerting.
Automation & Tooling:
Develop automation for failover, capacity management, and self-healing mechanisms to reduce operational risk.
Collaboration:
Partner with development, infrastructure, and production support teams to embed reliability into the SDLC.
Continuous Improvement:
Analyze service risk assessments and production incidents to identify systemic issues and drive long-term improvements.
Culture Building:
Promote operational excellence and a mindset of designing for failure across all engineering teams.
Required Skills & Experience:
Technical Expertise:
12+ years in softwar
📌 Principle Sre Pune
🏢 Barclays
📍 Pune