Role Overview:
We are looking for a Tech Ops / SRE Manager to own production stability, incident response, monitoring, uptime, root cause analysis, and reliability operations across our SaaS platforms. This role will ensure production issues are handled quickly, recurring failures are reduced, and reliability standards are implemented across AWS and GCP environments.
Key Responsibilities:
- Own production support and reliability operations for SaaS platforms.
- Define, monitor, and improve uptime, SLAs, SLOs, SLIs, MTTR, and incident trends.
- Manage incident response, escalation, communication, and post-incident reviews.
- Ensure proper root cause analysis and corrective action tracking.
- Build and maintain monitoring, alerting, and operational dashboards.
- Coordinate with Engineering, Cloud, InfoSec, and Product teams during incidents and releases.
- Create and maintain runbooks, escalation matrices, and production support processes.
- Identify recurring production issues and ensure permanent fixes are implemented.
- Support disaster recovery testing, backup validation, and business continuity readiness.
- Improve system reliability through automation, health checks, and preventive controls.
Required Experience:
- 8-12 years of experience in Tech Ops, Production Support, Infrastructure, SRE, or Platform Operations.
- 3 5 years in a lead or manager role.
- Mandatory experience supporting SaaS or cloud-based production systems.
- Strong experience with incident management, RCA, monitoring, alerting, and production operations.
- Experience with AWS, GCP, Kubernetes, Docker, and modern observability tools.
- Prior experience in fintech, payments, banking, compliance SaaS, or regulated environments is strongly preferred.
Ideal Candidate:
A calm and structured production leader with SaaS operations experience and robust preference for fintech/payments/banking exposure. The candidate should be able to manage incidents, improve uptime, reduce recurring issues, and bring discipline to production operations.
Mandatory Experience Requirement
Mandatory: Prior experience managing production reliability for a SaaS or cloud-based product company.
Strongly preferred: Experience in fintech, payments, banking, or regulated SaaS environments.
5+ years of experience ,
3 5 years in a lead or manager role
Disclaimer : This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Site Reliability Engineer (SRE) Engineer (India)
🏢 Flexm
📍 India