30 Sep
|
Source right
|
Bengaluru
30 Sep
Source right
Bengaluru
Position: Site Reliability Engineer (CE58ST RM 4459)
Key responsibilities
Production operations & reliability
- Own and operate production-grade systems in a 24×7 follow-the-sun on-call rotation, ensuring SLAs, SLOs, and error budgets are consistently met.
- Lead incident response, triage, and resolution for P1/P2 incidents, driving clear communication across engineering and business stakeholders throughout.
- Conduct thorough root cause analysis (RCA) and post-incident reviews (PIRs); drive follow-through on action items to prevent recurrence.
- Proactively identify reliability risks, capacity constraints, and performance bottlenecks before they become incidents.
- Define and maintain runbooks, escalation procedures, and operational playbooks for all critical services.
- Infrastructure & automation
- Design and manage secure, scalable Azure/AWS infrastructure using Terraform.
- Operate and optimize Kubernetes environments on EKS/AKS, including networking, scaling, monitoring, and reliability.
- Build and maintain CI/CD and GitOps delivery pipelines with GitLab CI and ArgoCD.
- Automate operational tasks through scripting and platform tooling to reduce manual effort and improve consistency.
- On-call & incident management tooling
- Administer and optimise on-call management platforms such as PagerDuty or Zenduty — including escalation policies, alert routing, and on-call scheduling across time zones.
- Ensure reliability and observability through effective use of ELK, Prometheus, Grafana, Datadog, or equivalent tooling, keeping alerts actionable and reducing noise.
- Lead incident response, support the on-call rotation, and champion a culture of blameless incident management and continuous improvement across the SRE team.
- Cloud platform ownership
- Own Azure/AWS cloud infrastructure, including networking, IAM, security posture, cost awareness, and service reliability.
- Support secure cloud environments aligned with SOC2 and ISO27001 standards.
- Contribute to cloud architecture and deployment standards with a solid focus on operability and resilience.
- Collaboration & documentation
- Partner with engineering, security, and product teams to improve reliability and operational readiness.
- Support a global setup across multiple time zones and collaborate effectively with international teams.
- Create runbooks, incident reports, and technical documentation for operational excellence.
- Mentor team members and contribute to continuous improvement across the SRE function.
Required Qualifications
- 5–8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles.
- Strong hands-on experience with Azure/AWS, Terraform, Kubernetes, and ArgoCD in production environments.
- Experience operating EKS-based platforms, including networking, scaling, monitoring, and troubleshooting.
- Strong knowledge of CI/CD, GitOps, and automation practices, with hands-on use of GitLab CI and ArgoCD.
- Experience managing production systems in a 24×7 environment, including incident response and on-call practices.
- Solid Linux and cloud networking background.
- Knowledge of managing cloud environments aligned with SOC2 and ISO27001 standards.
- Experience with observability tools such as ELK and Prometheus.
- Strong scripting skills in Python, Bash, or Go.
- Strong written and verbal communication skills, with the ability to work effectively across global teams and time zones.
- Familiarity with security integration in CI/CD pipelines, including SAST, DAST, container scanning, and secrets management.
Preferred Qualifications
- Engineering degree in computer science or equivalent.
- Cloud certifications such as AWS/Azure SysOps / DevOps Engineer, or CKA/CKAD.
- Exposure to ITSM or change management processes in regulated industries (healthcare, fintech, or similar).
📌 Site Reliability EngineerST RM 4459) (Bengaluru)
🏢 Source right
📍 Bengaluru