06 Sep
|
Consultbae India
|
India
06 Sep
Consultbae India
India
Role Overview
We are looking for an experienced SRE Engineer to ensure the reliability, scalability, and performance of large-scale distributed systems in a production environment.
Key Responsibilities
- Manage and optimize production environments across AWS, Kubernetes, Docker, and Linux/Unix.
- Build and maintain CI/CD pipelines and Infrastructure as Code using Terraform and Ansible.
- Implement observability and monitoring using Prometheus, Grafana, OpenTelemetry, or Datadog.
- Set up SLO-based alerting and incident detection mechanisms.
- Troubleshoot production issues and participate in incident response.
- Improve system reliability, availability, scalability, and performance.
- Participate in the production on-call rotation,
including some weekends approximately once every 2–3 weeks.
Must-Have Skills
- 4–6 years of experience in SRE, DevOps, or Production Engineering.
- Robust hands-on experience with AWS, Kubernetes, Terraform, Grafana, and Prometheus.
- Strong knowledge of Linux/Unix and Docker.
- Experience with CI/CD and Infrastructure as Code.
- Hands-on experience with observability, monitoring, alerting, and SLOs.
- Experience working with large-scale distributed systems in production.
📌 SRE Engineer (India)
🏢 Consultbae India
📍 India