Senior Consultant | Site Reliability Engineering | Bengaluru
Role Overview
We are looking for a highly skilled Site Reliability Engineer (SRE) to manage and scale mission-critical, production-grade distributed systems running on AWS. The ideal candidate will focus on reliability, automation, observability, and operational excellence while minimizing toil and improving system availability. Maintaining and improving 4 Nines of uptime to 5 Nines with engineering efforts.
This role requires deep technical expertise in cloud-native technologies, Kubernetes, infrastructure automation, Linux administration and TCP/IP fundamentals, and solid troubleshooting capabilities for distributed systems. The candidate needs to participate in the overall lifecycle management of mission critical banking services with a 24x7 operations mode in an rotational on-call basis. The job requires the candidate to have strong troubleshooting skills in a distributed environment spread across multiple cloud environments. The bare minimum ask would be to maintain high level of agility,
learnability and adaptability in different scenarios. An engineer with a zeal to learn fast and having a bias for action would be the best fit for the role.
Key Responsibilities
Reliability & Operations
- Own end-to-end production systems reliability, availability, scalability, cost and performance.
- Troubleshooting EKS from SRE perspective
- Route53, Connectivity, Networking, RDS, S3
- Participate in 24x7 on-call rotations and handle high-severity incidents and document the learnings on ongoing basis.
- Establish and manage SLI, SLO, SLA, Error Budgets, and operational metrics for mission critical services and partner with engineering teams with full accountability for upholding the SLOs.
- Partner with the various engineering, operations and cloud management teams to deliver highly reliable service in a timely manner.
Cloud & Infrastructure
- Design, deploy, and manage infrastructure on Google Cloud Platfor