We are looking for an experienced SRE / Production Engineering skilled who can drive reliability, automation, observability, and production stability across a cloud-based environment.
? Key Responsibilities
- Build and promote an SRE culture by sharing best practices, documentation, and automation across engineering teams.
- Automate manual processes using software, scripting, and infrastructure automation .
- Troubleshoot complex OS, networking, database, and cloud infrastructure issues.
- Handle live production incidents , perform RCA, and implement permanent solutions.
- Monitor and improve application performance, availability, stability, and reliability .
- Design and implement solutions to enhance observability and operational efficiency .
- Manage deployment and orchestration of servers, Docker containers, databases, and backend infrastructure .
- Develop Runbooks / SOPs for recurring production issues.
- Conduct regular Incident Analysis and Problem Management to prevent recurring incidents.
? Mandatory Skills ✅ Site Reliability Engineering / Production Engineering
✅ Kubernetes & Docker
✅ Terraform / CloudFormation / Ansible
✅ AWS / Azure / Cloud technologies
✅ Kafka / Confluent Platform Kafka
✅ Infrastructure & Application Monitoring
✅ Linux / OS troubleshooting
✅ Networking & Database troubleshooting
✅ Production Incident Management & RCA
✅ Distributed Technologies
? Education
B.Tech / BE in Computer Science or a related discipline preferred.
? Interested candidates can share their updated CV below link