17 Sep
|
Kestra
|
Bengaluru
What youll Do:
- Lead and develop a high-performing team of Dev Ops and Site Reliability Engineers through coaching, mentoring, performance management, and career development.
- Own the Dev Ops and SRE strategy, roadmap, and operational priorities for the Advisor Platform.
- Drive platform reliability, availability, scalability, resiliency, and operational excellence across production environments.
- Establish and track key reliability and operational metrics, including SLIs, SLOs, error budgets, availability, MTTR, deployment frequency, and service health indicators.
- Lead the implementation and continuous improvement of CI/CD pipelines, Infrastructure as Code, automation frameworks, and cloud-native engineering practices.
- Champion automation initiatives that reduce operational toil and improve platform stability, deployment efficiency, and engineering productivity.
- Ensure robust observability through monitoring, logging, tracing, alerting, and operational dashboards that provide actionable insights.
- Lead major incident management activities, including stakeholder communications, incident command, root cause analysis, postmortems, and corrective action tracking.
- Partner closely with Software Engineering teams to embed reliability, security, observability, and operational readiness throughout the software development lifecycle.
- Collaborate with Infrastructure, Cloud, Security, Database, and Network teams to design, operate, and optimize highly available and secure platforms.
- Drive capacity planning, performance engineering, scalability initiatives, and disaster recovery preparedness to support business growth.
- Establish and continuously improve operational processes, including change management,
release governance, production readiness reviews, problem management, and knowledge management.
- Promote reliability-first engineering principles and operational ownership across Engineering and Technology teams.
- Support cloud modernization and platform transformation initiatives while maintaining operational stability and governance.
- Communicate platform health, service performance, operational risks, incident trends, and improvement initiatives to senior leadership through clear, data-driven reporting.
- Ensure compliance with security, audit, regulatory, data protection, risk management, and operational governance requirements.
- Foster a culture of accountability, collaboration, continuous learning, innovation, and continuous improvement.
What You Bring:
- Bachelor's degree in Engineering, Computer Science, Information Technology, or a related technical discipline. Engineering degree preferred.
- 8+ years of experience in Dev Ops, Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, Cloud Engineering, or related disciplines.
- 3+ years of experience leading Dev Ops, SRE, Platform Engineering, Infrastructure, or Production Operations teams.
- Solid understanding of Dev Ops and SRE principles, including CI/CD, Infrastructure as Code, automation, observability, incident management,
SLIs, SLOs, and error budgets.
- Hands-on experience with cloud platforms such as AWS, Azure, or GCP.
- Experience building and operating highly available, scalable, resilient enterprise applications and platforms.
- Strong expertise with monitoring, logging, tracing, alerting, and observability tools.
- Experience with containerization and orchestration technologies such as Docker and Kubernetes.
- Experience implementing Infrastructure as Code and automation technologies such as Terraform, Ansible, or equivalent tools.
- Experience with modern CI/CD platforms such as Git Hub Actions, Azure Dev Ops, Jenkins, Git Lab CI, or similar technologies.
- Proven track record of leading major incident response, problem management, and operational excellence initiatives.
- Experience with disaster recovery, business continuity, change management, and release governance processes.
- Strong analytical, troubleshooting, and problem-solving skills with the ability to manage high-severity production incidents.
- Excellent leadership, communication, stakeholder management, and cross-functional collaboration skills.
- Experience supporting financial services, wealth management, fintech, or other regulated industry platforms is preferred.
- Experience leading cloud modernization, Dev Ops transformation, platform engineering, or enterprise SRE initiatives is preferred.
- Familiarity with security, compliance, and governance frameworks such as SOC, SOX, PCI DSS, SOC 2, and ITIL is preferred.
- Master's degree in Engineering, Computer Science, Information Systems, or MBA with strong technical depth is a plus.
📌 Manager, DevOps & Site Reliability Engineering (SRE) (Bengaluru)
🏢 Kestra
📍 Bengaluru