FNP is looking for an experienced SRE Manager to lead our Site Reliability Engineering team. The ideal candidate will have a strong background in DevOps practices, system reliability, cloud infrastructure, automation, and team leadership.
Responsibilities
- Team Leadership: Lead, mentor, and manage a team of SRE and DevOps engineers while building a strong culture of ownership and reliability.
- SRE Best Practices: Define and implement SRE practices including SLIs, SLOs, and error budgets.
- System Reliability: Ensure the reliability, scalability, availability, and performance of critical systems.
- Automation: Drive automation initiatives to improve operational efficiency and reduce manual intervention.
- Cross-Functional Collaboration: Collaborate with engineering and other cross-functional teams to ensure reliable and scalable technology solutions.
- CI/CD Release Management: Own CI/CD pipelines and release management processes.
- Incident Management: Lead incident response, troubleshooting, and Root Cause Analysis (RCA) processes.
- Monitoring Observability: Establish and improve monitoring and observability frameworks across technology platforms.
- Cloud Infrastructure: Manage and optimize cloud infrastructure across AWS, Azure, or GCP environments.
- Disaster Recovery: Implement and maintain disaster recovery and business continuity plans.
Qualifications
Experience
- 7+ years of experience in SRE or DevOps roles.
- 3+ years of team management experience.
- Experience managing production-grade cloud infrastructure.
- Experience in e-commerce platforms is a good to have.
Required Skills
- Robust experience with AWS, Azure, GCP, or similar cloud platforms.
- Knowledge of CI/CD tools such as Jenkins and GitLab CI.
- Experience with Docker and Kubernetes.
- Scripting skills in Python and Bash.
- Knowledge of Terraform and/or CloudFormation.
- Experience with monitoring tools such as Prometheus, Grafana, and ELK.
Preferred Qualifications
- Experience working with microservices architectures.
- Cloud certifications are a plus.
- Strong problem-solving skills.
- Knowledge of chaos engineering is a good to have.
Key Competencies
- Leadership
- Communication
- Ownership
- Stakeholder Management
- Problem Solving
- System Reliability
- Team Management
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Engineering Manager - Site Reliability Engineer (Gurugram)
🏢 Ferns N Petals
📍 Gurugram
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.