01 Sep
|
Nium India
|
Bengaluru
01 Sep
Nium India
Bengaluru
The Role
Nium is looking for a Senior Manager, Site Reliability Engineering to lead the teams responsible for the availability, performance, scalability, and operational excellence of our global payments platform. This is a hands-on leadership role: you will build and grow a team of SREs, define the reliability roadmap, and partner closely with product engineering, security, and infrastructure teams to ensure Niums systems meet the always-on expectations of a regulated financial platform operating across 100+ markets.
You will own incident management, observability, capacity planning, and production readiness practices, while championing a culture of blameless postmortems, proactive risk reduction, and engineering-driven automation. This role sits at the intersection of engineering leadership and operational rigor, and reports into Niums engineering leadership.
Responsibilities
- Lead, mentor, and grow a team of SREs and reliability engineers across multiple time zones, setting explicit goals, career paths, and performance expectations.
- Own Niums reliability strategy - defining and driving SLIs/SLOs/error budgets across critical payment, card issuance, and compliance services.
- Drive incident management end-to-end: on-call structure, escalation paths, major incident response, and blameless postmortems that produce durable fixes, not just tickets.
- Partner with product engineering leaders to embed reliability, scalability, and operational readiness into the software development lifecycle from design through launch.
- Build and scale observability (metrics, logging, tracing, alerting) so that issues are detected and diagnosed before they impact customers or partner banks.
- Lead capacity planning and performance engineering for systems processing high-volume, real-time financial transactions across a global, multi-region infrastructure.
- Champion automation and self-healing systems to reduce toil, eliminate manual runbooks,
and improve mean-time-to-detect and mean-time-to-resolve.
- Own disaster recovery, business continuity, and chaos engineering practices, running regular game days to validate resilience assumptions.
- Collaborate with Security and Compliance teams to ensure infrastructure practices meet regulatory and audit requirements (PCI-DSS, SOC 2, ISO 27001, and regional financial regulations).
- Manage the reliability budget: tooling investments, cloud cost/performance trade-offs, and staffing plans, in partnership with finance and engineering leadership.
- Represent SRE in executive reviews, translating technical risk and system health into business-relevant reporting for leadership and the board.
- Establish and continuously refine production readiness reviews, runbooks, and operational standards across all engineering teams.
Requirements
- 10+ years of experience in software engineering, infrastructure, or site reliability engineering, with 4+ years in a people-management or technical leadership role leading SRE/DevOps/Infrastructure teams.
- Proven track record operating and scaling production systems for a high-availability, transaction-heavy platform - fintech, payments, banking, or e-commerce experience strongly preferred.
- Deep hands-on expertise with cloud infrastructure (AWS) Kubernetes, container orchestration, and infrastructure-as-code (Terraform, CloudFormation, or similar).
- Strong background in observability stacks (Prometheus, Grafana, Datadog, ELK/OpenSearch, or equivalent) and building alerting that reduces noise while catching real issues.
- Demonstrated experience defining and operationalizing SLOs/error budgets, and using them to drive engineering prioritization.
Disclaimer: This job posting and Location has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Sr Manager, Site Reliability Engineering (SRE) (Bengaluru)
🏢 Nium India
📍 Bengaluru