24 Sep
|
Recro
|
Bengaluru
Site Reliability Engineer
Experience:
3+ years
Location:
Bangalore
Shift:
Rotational shifts
Role Overview
We are looking for an SRE to ensure the reliability, scalability, availability, and observability of high-traffic production systems
. The role involves infrastructure automation, Kubernetes operations, monitoring, incident management, and continuous improvement of production reliability.
Key Responsibilities
- Manage and troubleshoot
AWS/GCP infrastructure and Kubernetes environments in production
.
- Work with
Kubernetes, Redis, Kafka, Solr/Elasticsearch and related infrastructure components.
- Build and maintain
CI/CD pipelines, Terraform/Helm-based infrastructure, and automation
.
- Implement monitoring and alerting using
Prometheus, Grafana, ELK/Loki and distributed tracing tools.
- Participate in rotational on-call and shifts
, handling P1/P2 incidents, troubleshooting, RCA, and post-incident reviews.
- Develop
SLIs/SLOs, alerts, dashboards, runbooks
, and reliability improvements.
- Automate repetitive operational tasks using
Python, Shell/Bash, or Go to reduce manual toil.
- Work on capacity planning, autoscaling, performance optimization, security patching, and cost optimization
.
- Troubleshoot
Linux, networking, DNS, TCP/IP, load balancing, TLS/HTTPS
, and application/infrastructure issues.
- Drive preventive actions through RCA, automation, self-healing, and improved deployment/recovery processes.
Must-Have Skills
- 3+ years of experience in
SRE / DevOps / Infrastructure Engineering
- Experience with high-traffic or large-scale production environments
- Hands-on
Kubernetes in production
- Strong experience with
AWS or GCP
- Terraform and Infrastructure as Code (
IaC
)
- Docker and
CI/CD
- Prometheus & Grafana
- Linux administration and troubleshooting
- Python / Bash / Shell scripting
- Production incident management, RCA and on-call experience
- Valuable understanding of networking fundamentals
Good to Have
- Helm, ArgoCD/GitOps
- ELK/Loki
- Redis, Kafka, Solr/Elasticsearch
- SLI/SLO, SLA and error-budget concepts
- Ansible
- Distributed tracing / OpenTelemetry
Note:
This is a rotational-shift/on-call role
, so candidates should be comfortable supporting production systems across different shifts.
📌 Site Reliability Engineer (Bengaluru)
🏢 Recro
📍 Bengaluru