Role &
Responsibilities :
- Own the reliability, availability, scalability, and performance of production systems and critical services.
- Design, implement, and maintain highly available cloud infrastructure across AWS/Azure/GCP environments.
- Define and improve SLOs, SLIs, and SLAs to measure and enhance system reliability.
- Lead incident response, troubleshoot complex production issues, perform root cause analysis (RCA), and drive preventive actions.
- Build and maintain automation for infrastructure provisioning, deployments, monitoring, and operational workflows.
- Manage and optimize Kubernetes clusters, containerized workloads, and cloud-native architectures.
- Develop and maintain CI/CD pipelines to enable safe, frequent, and reliable software releases.
- Implement observability solutions using tools such as Prometheus, Grafana, Datadog, New Relic, Splunk, or ELK.
- Improve system performance through capacity planning, load testing, scalability improvements, and tuning.
- Establish disaster recovery strategies, backup processes, and business continuity plans.
- Automate repetitive operational tasks using scripting and infrastructure-as-code practices.
- Collaborate with engineering, security, and product teams to improve system resilience.
- Participate in on-call rotations and provide technical leadership during high-priority incidents.
Preferred Candidate Profile :
- Bachelors degree in Computer Science, Information Technology, Engineering, or a related field.
- 5-10 years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or Production Engineering roles.
- Hands-on expertise with at least one major cloud platform:
- AWS
2.
Microsoft
Azure
3.
Google Cloud
Platform (GCP)
- Robust knowledge of:
- Linux administration
- Networking fundamentals (TCP/IP, DNS, HTTP, load balancing)
- Distributed systems concepts
- System troubleshooting and performance optimization
- Proficiency in scripting/programming languages:
- Python
- Bash
- Go (preferred)
- Experience with monitoring and observability platforms:
- Prometheus
- Grafana
- Datadog
- Splunk
- ELK Stack
Strong understanding of
- Incident management
- Root cause analysis
- Error budgets
- Capacity planning
- Reliability metrics
- Familiarity with security best practices, IAM, vulnerability management, and compliance requirements.
- Excellent communication skills with the ability to work with globally distributed teams.
- Ability to independently manage production responsibilities in a remote environment.
📌 Senior Reliability Engineer - Cloud Infrastructure (India)
🏢 Good
📍 India