Job Summary
SRE and DevOps Team Engineer | 100% Remote - India
We are seeking an experienced SRE DevOps Manager to lead the team responsible for the reliability, automation, and scale of our SaaS platform. This is a hands-on role: you will work closely with engineering leadership to improve processes, operate incident management and on-call practices, and guide and develop a team of engineers - while remaining close enough to the systems to support the team through difficult operational challenges with credibility.
You will be accountable for the availability, performance, and operational maturity of our cloud environments, and for promoting a team culture grounded in reliability, ownership, and continuous improvement. The role coordinates day-to-day across Indian Standard Time (IST) and US time zones.
Responsibilities
- Team leadership: Lead, mentor, and grow a team of SRE/DevOps engineers. Work closely with engineering leadership to assess team needs and develop talent.
- Incident management: Oversee the incident management process end to end - on-call rotations, escalation paths, incident command, postmortems, and follow-through on root cause analysis (RCA) and preventative actions.
- Reliability practices: Define and drive SRE principles including SLIs, SLOs, error budgets, capacity planning, and observability standards across services. Champion a culture of reliability and operational excellence.
- Process improvement: Partner with engineering leadership to identify and implement process improvements across the team s operational priorities, balancing feature delivery, reliability work, security, and cost efficiency. Guide and support the team s day-to-day work.
- Cross-time-zone collaboration: Coordinate a distributed team and support model across IST and US time zones, ensuring effective handoffs, coverage, and communication.
- Leadership partnership:
Interface with engineering leadership to align infrastructure investments with business and compliance goals. Report on team health, reliability metrics, and operational performance. Assist engineering leadership in managing the team, including department-wide initiatives and roadmap planning.
Required Experience & Skills
- Leadership experience: 2-3 years of experience as a team lead or manager in a DevOps/SRE environment.
- Distributed teams: Proven experience leading and coordinating distributed teams across time zones, specifically IST and US, including on-call coverage and handoffs.
- Incident management: Experience running an incident management process - on-call, escalation, incident command, severity frameworks, and driving blameless postmortems and RCA to completion.
- Cloud infrastructure: Strong hands-on background across multi-cloud environments (e.g., AWS, Kubernetes - GKE EKS, managed databases, object storage, networking, IAM) with flexibility to work across multiple cloud providers.
- Platform Engineering: Solid experience with Infrastructure-as-Code practices (Terraform, Pulumi, or similar) and CI/CD tooling (GitHub Actions, Jenkins, or similar).
- Observability : Experience with modern observability practices and tooling (metrics, logging, tracing, alerting).
Preferred Qualifications
- Experience improving and scaling an SRE/DevOps function, on-call program, or reliability practice.
- Experience in FinTech or another regulated setting with security and compliance requirements.
- Experience optimizing data infrastructure and database operations.
- Experience improving observability across multiple cloud technologies.
- A background in designing for cost-efficiency, resilience, and security in cloud infrastructure.
Disclaimer: This job posting & Location has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 SRE and Devops Team Manager (India)
🏢 Eltropy
📍 India