We are looking for an experienced SRE / Production Engineering qualified who can drive reliability, automation, observability, and production stability across a cloud-based environment.
? Key Responsibilities
- Build and promote an SRE culture by sharing best practices, documentation, and automation across engineering teams.
- Automate manual processes using software, scripting, and infrastructure automation.
- Troubleshoot complex OS, networking, database, and cloud infrastructure issues.
- Handle live production incidents, perform RCA, and implement permanent solutions.
- Monitor and improve application performance, availability, stability, and reliability.
- Design and implement solutions to enhance observability and operational efficiency.
- Manage deployment and orchestration of servers, Docker containers, databases, and backend infrastructure.
- Develop Runbooks / SOPs for recurring production issues.
- Conduct regular Incident Analysis and Problem Management to prevent recurring incidents.
? Mandatory Skills ✅ Site Reliability Engineering / Production Engineering
✅ Kubernetes & Docker
✅ Terraform / CloudFormation / Ansible
✅ AWS / Azure / Cloud technologies
✅ Kafka / Confluent Platform Kafka
✅ Infrastructure & Application Monitoring
✅ Linux / OS troubleshooting
✅ Networking & Database troubleshooting
✅ Production Incident Management & RCA
✅ Distributed Technologies
? Education
B.Tech / BE in Computer Science or a related discipline preferred.
? Interested candidates can share their updated CV below link