- Automate manual processes and improve operational efficiency
- Troubleshoot complex OS, Networking & Database issues in a cloud-based SaaS environment
- Handle live production incidents and drive effective resolution
- Improve system performance, availability, reliability & scalability
- Design and develop solutions to enhance observability and product reliability
- Create and maintain Runbooks & SOPs for recurring production issues
- Conduct Incident Analysis/RCA and implement permanent fixes
- Drive DevOps and automation initiatives across infrastructure and applications
Preferred candidate profile
- Hands-on experience with AWS & Cloud Infrastructure
- Docker & Kubernetes
- Strong programming skills in Python
- Terraform, CloudFormation & Ansible
- Solid Linux Internals & TCP/IP Networking troubleshooting
- Experience with Kafka & Confluent Kafka Platform
- Infrastructure & application monitoring/performance analysis
- Strong DevOps, Network & Infrastructure expertise
- Experience working with large-scale databases and distributed systems
Education: B.Tech in Computer Science or related discipline preferred.