04 Aug
|
Falabella India
|
Bengaluru
04 Aug
Falabella India
Bengaluru
We are looking for an experienced Senior Site Reliability Engineer (SRE) to join our platform engineering team. The ideal candidate will be responsible for designing, building, and operating highly available, scalable, secure, and cost-efficient cloud infrastructure platforms. This role requires strong expertise in Kubernetes, cloud platforms, observability, automation, incident management, and reliability engineering practices.
The Senior SRE will collaborate closely with software engineering, security, infrastructure, and operations teams to improve system reliability, performance, scalability, and developer productivity.
Reliability Engineering The core responsibilities for the job include the following:
- Design, implement, and maintain highly available and resilient production systems.
- Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
- Drive reliability improvements through automation and engineering best practices.
- Perform capacity planning and performance optimisation for critical systems.
- Conduct failure analysis and implement preventive measures.
Cloud Infrastructure And Kubernetes
- Design, deploy, and manage large-scale Kubernetes environments.
- Manage cloud infrastructure across AWS, GCP, or Azure environments.
- Implement Infrastructure as Code (IaC) using Terraform.
- Manage service mesh platforms such as Istio.
- Optimise infrastructure utilisation and cloud costs.
Observability And Monitoring
- Build and maintain monitoring, logging, and alerting platforms.
- Implement observability solutions using tools such as Prometheus, Grafana, Datadog, VictoriaMetrics, and Loki.
- Develop dashboards, alerts, and automated operational workflows.
- Monitor system performance, availability, latency, and business metrics.
Incident Management And Operations
- Lead production incident response and troubleshooting activities.
- Participate in on-call rotations and major incident management.
- Conduct root cause analysis (RCA) and post-incident reviews.
- Drive continuous improvements to reduce operational toil.
- Establish operational runbooks and automation frameworks.
Automation And DevOps
- Develop automation scripts and tooling using Python, Go, or Bash.
- Implement CI/CD pipelines and GitOps workflows.
- Build self-service infrastructure capabilities for engineering teams.
- Automate operational tasks and infrastructure provisioning.
Security And Compliance
- Collaborate with security teams to implement DevSecOps practices.
- Ensure infrastructure compliance with organisational standards.
- Implement secure access controls, secrets management, and audit processes.
- Support vulnerability management and remediation efforts.
FinOps And Optimisation
- Monitor cloud spending and identify optimisation opportunities.
- Implement rightsizing, reserved instance, and committed-use strategies.
- Analyse infrastructure costs and recommend cost-saving initiatives.
- Partner with engineering teams to improve resource efficiency.
Requirements
- Bachelor's degree in computer science, engineering, or a related field.
- 4+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering.
- Strong experience managing production Kubernetes environments.
- Hands-on expertise with public cloud platforms (AWS, GCP, or Azure).
- Strong knowledge of Linux systems administration and networking.
- Experience with Infrastructure as Code tools such as Terraform.
- Expertise in observability and monitoring platforms.
- Experience implementing CI/CD and GitOps practices.
- Strong scripting and programming skills (Python, Go, Bash).
- Experience with incident management and root cause analysis.
Preferred Qualifications
- Experience operating large-scale distributed systems.
- Experience with service mesh technologies (Istio, Linkerd).
- Knowledge of database reliability engineering (PostgreSQL, MySQL, Redis).
- Experience with security and compliance frameworks.
- FinOps and cloud cost optimisation experience.
- Experience managing multi-region and multi-cloud environments.
- Kubernetes certifications (CKA, CKAD, CKS) are preferred.
Technical Skills
- Cloud Platforms: AWS, Google Cloud Platform (GCP), Microsoft Azure.
- Container and Platform: Kubernetes, Docker, Helm, Istio.
- Infrastructure as Code: Terraform, Ansible.
- Observability: Prometheus, Grafana, Datadog, Elasticsearch, VictoriaMetrics, Loki.
- CI/CD and GitOps: GitHub Actions, GitLab CI, Jenkins, ArgoCD.
- Databases: PostgreSQL, MySQL, Redis.
- Programming: Python, Go, Bash.
Soft Skills
- Strong troubleshooting and analytical skills.
- Excellent communication and stakeholder management abilities.
- Ability to lead complex technical initiatives.
- Robust ownership mindset and operational excellence.
- Ability to mentor engineers and drive engineering best practice.
This job was posted by Priyanka R N from Falabella.
📌 Senior Site Reliability Engineer (Bengaluru)
🏢 Falabella India
📍 Bengaluru