Role Overview :
As a Lead DevOps / Site Reliability Engineer, you will serve as the architectural backbone for our high-scale production environments, ensuring that our platforms remain resilient, performant, and secure. You will work closely with cross-functional engineering teams, product managers, and infrastructure stakeholders to bridge the gap between development and operations.
Your daily efforts will directly influence the end-user experience by minimizing downtime, optimizing cloud resource utilization, and automating complex deployment pipelines. By championing SRE best practices, you will drive the technical strategy that allows our business to scale rapidly while maintaining the highest standards of system reliability and operational excellence.
Key Responsibilities :
- Architect and maintain robust, scalable cloud infrastructure to support high-traffic applications, ensuring 99.99% availability for our global customer base.
- Implement advanced CI/CD pipelines to accelerate release cycles, enabling developers to ship features faster with automated quality and security gates.
- Lead incident response and post-mortem analysis to identify root causes of system failures, proactively preventing recurring issues and improving overall platform stability.
- Optimize cloud infrastructure costs by implementing intelligent resource management and auto-scaling strategies, directly contributing to the company's bottom-line efficiency.
- Mentor junior engineers and foster a culture of reliability, sharing technical expertise to elevate the collective engineering standards of the team.
Required Skillset :
- Demonstrated expertise in designing and managing complex distributed systems on AWS, Azure, or GCP, with a deep understanding of container orchestration using Kubernetes.
- Proven ability to write clean, maintainable infrastructure-as-code using Terraform or CloudFormation to manage multi-region environments.
- Strong proficiency in scripting and automation using Python, Go, or Bash to eliminate manual toil and streamline operational workflows.
- Exceptional communication skills, with the ability to articulate complex technical trade-offs to non-technical stakeholders and lead cross-departmental initiatives.
- A collaborative mindset that thrives in a hybrid work workplace, balancing independent problem-solving with active participation in team-wide architectural reviews.
- A solid academic foundation in Computer Science or a related engineering discipline, complemented by a track record of continuous learning in the evolving DevOps landscape.
- Candidates must possess 6 to 14 years of relevant experience in managing large-scale production environments, with a preference for those who have navigated high-growth startup or enterprise-scale transitions.
📌 Lead DevOps/Site Reliability Engineer (India)
🏢 HyrMe
📍 India