Experience: 7 to 10 Years
Location: Pune, Kharadi
Work Mode: 5 Days Work from Office
Joining: Immediate Joiners Only
We are looking for a highly skilled Senior/Lead DevOps Engineer with strong Site Reliability Engineering (SRE) experience to join our growing engineering team.
The ideal candidate should have strong hands-on expertise in AWS, Kubernetes, Infrastructure as Code, CI/CD automation, and production reliability engineering .
This role requires an engineer who can design, automate, deploy, monitor, troubleshoot, and optimize cloud-native environments while ensuring high availability, scalability, security, and operational excellence .
Key Responsibilities
- Design, deploy, and manage highly available cloud infrastructure on AWS.
- Build, manage, and troubleshoot production-grade Kubernetes environments.
- Implement Infrastructure as Code using Terraform and AWS CloudFormation .
- Develop, maintain, and optimize CI/CD pipelines using Jenkins .
- Drive automation across infrastructure provisioning, deployments, monitoring, and incident management.
- Work closely with engineering teams to improve system reliability, performance, and scalability.
- Handle production incidents, perform root cause analysis, and drive post-incident improvements.
- Implement and maintain monitoring, alerting, logging, and observability solutions.
- Optimize cloud infrastructure for cost, security, performance, and reliability.
- Support containerized applications and microservices architectures.
- Maintain and improve deployment workflows using Git and Bitbucket .
- Contribute to SRE best practices, including SLIs, SLOs, and Error Budgets .
Mandatory Skills
- 7 to 10 years of hands-on DevOps / SRE experience .
- Strong hands-on experience with AWS .
- Production-grade Kubernetes (K8s) experience.
- Strong expertise in Terraform .
- Hands-on experience with AWS CloudFormation .
- Strong experience creating and managing Jenkins pipelines .
- Strong understanding of CI/CD implementation and automation .
- Strong Linux administration and troubleshooting skills.
- Proficiency in Shell Scripting, Bash, or Python .
- Experience with monitoring and logging tools such as Prometheus, Grafana, ELK, CloudWatch , or similar tools.
- Experience managing production environments and handling critical incidents.
Preferred Skills
- Robust Site Reliability Engineering experience.
- Bitbucket administration and pipeline integration.
- Docker and containerization technologies.
- Exposure to Azure .
- Security, compliance, and infrastructure governance.
- Experience working in Agile environments.
Ideal Candidate Profile
- Strong hands-on engineer who enjoys solving complex infrastructure and reliability challenges.
- Individual contributor, not a people-management role .
- Experience supporting large-scale production environments.
- Strong understanding of reliability engineering and operational excellence.
- Startup or product-based company experience is preferred.
- Excellent troubleshooting and problem-solving skills.
- Strong communication and stakeholder management skills.
- Comfortable taking end-to-end ownership of infrastructure, deployments, and production reliability.
Nice to Have
- AWS certifications.
- Kubernetes certifications.
- Azure exposure.
- Experience in healthcare, fintech, SaaS, or other product-based organizations.
📌 Site Reliability Engineer Lead (Immediate Joiner) (Pune)
🏢 HiLabs
📍 Pune