Director – Observability, Response & Reliability (ORR) (Noida)

Director – Observability, Response & Reliability (ORR) (Noida)

31 Jul
|
NetSpend
|
Noida

31 Jul

NetSpend

Noida

About The Company Netspend Corporation is a global, vertically-integrated financial services and technology company dedicated to the delivery of cutting-edge financial empowerment solutions to consumers worldwide. Netspend's financial products and services span prepaid, debit, cross-border payments, and loyalty solutions for consumers and enterprise partners. Netspend provides prepaid and debit account solutions that connect customers with secure, convenient access to global payment networks so they can manage their money and make everyday purchases.

With a nationwide U.S. retail network, customers can purchase and reload Netspend products at 130,000 reload points and over 100,000 distributing locations. Since our founding in 1999 by industry pioneers, Netspend products have processed billions of dollars in transaction volume and served millions of customers worldwide. The company is headquartered in Austin, Texas with employees worldwide.

Role Overview We are seeking a visionary Director – Observability, Response & Reliability (ORR) with 15+ years of overall technical experience, including 5+ years in engineering leadership, to serve as our primary authority on system resilience, full-stack observability, and enterprise incident management across Netspend's global financial ecosystem. In this high-visibility role, you will lead our India-based ORR organization, driving the strategy that transforms how Netspend monitors, predicts, responds to, and resolves critical operational events. You will establish a world-class Observability and AIOps practice, institutionalize Site Reliability Engineering (SRE) principles across all product teams, and safeguard systems handling ACH processing, core payment rails, millions of daily card transactions, and bank-sensitive data.

Key Responsibilities Observability Strategy & Telemetry Architecture

Unified Telemetry Vision: Design and execute the enterprise Observability roadmap, establishing full-stack, end-to-end visibility across microservices, legacy monoliths, cloud infrastructure, and data pipelines (Kafka, Cassandra).

MELT Standardizer: Define unified logging, metrics, traces, and synthetics standards (OpenTelemetry, Prometheus, Splunk, Dynatrace) across all engineering groups to enable granular transaction-level tracing for payment flows.

AIOps & Self-Healing Systems: Pioneer the adoption of AIOps, Machine Learning, and anomaly detection to shift the engineering culture from reactive alert firefighting to proactive noise reduction, predictive fault detection, and automated self-healing workflows.





Financial Data Observability: Build custom monitoring, alert thresholds, and real-time dashboards for critical FinTech protocols, payment rails, ACH file watching, and bank connectivity channels (Axway, SFTP, AS2).

Reliability Engineering, Resilience & SLO Management

SRE Practice Ownership: Institutionalize SRE paradigms (SLIs, SLOs, Error Budgets, Reliability Reviews) across all software product squads, embedding reliability directly into the software development lifecycle (SDLC).

Chaos Engineering & Testing: Establish proactive resilience practices, including failure injection, Chaos Engineering, and regular multi-AZ/multi-region Disaster Recovery (DR) simulations for critical financial services.

Cost-Effective Reliability (FinOps): Partner with cloud and finance leadership to balance extreme uptime demands with cost efficiency across Splunk, Dynatrace, AWS telemetry storage, and log ingestion limits.

Enterprise Incident Response & Operational Governance

Major Incident Command (P0/P1): Oversee the high-severity Incident Management framework, ensuring 24/7 incident readiness, rapid mean time to detect (MTTD), and swift mean time to resolve/recover (MTTR).

Blameless Post-Mortems & RCA: Champion a culture of psychological safety through rigorous, blameless post-mortems and Root Cause Analyses (RCAs) to drive systemic platform fixes and prevent recurring outages.

Executive Reliability Reporting: Establish transparent reporting frameworks to translate uptime metrics, availability SLAs, error budget consumption, and platform risks directly to C-suite leadership (CTO, CIO).

Strategic Cloud Transformation & Infrastructure Evolution

Strangler Pattern Execution: Provide reliability oversight for moving critical functions off legacy platforms (Xymon, Puppet, SVN) to cloud-native AWS architectures without risking live financial traffic or data loss.

CI/CD & IaC Standards: Enforce strict Infrastructure-as-Code (Terraform/Ansible) and GitLab CI/CD pipeline reliability, integrating automated canary deployments, security scans, and instant rollback mechanisms.

Security & Compliance Oversight: Partner with security teams to ensure all observability logging, tracing,



and automation frameworks comply strictly with PCI-DSS, SOC 2, mTLS, and banking audit standards.

Organizational Leadership & Vendor Ownership

India ORR Center of Excellence: Build, mentor, and scale a high-performing team of Site Reliability Engineers, Observability Specialists, and Incident Response Leads in India.

Talent Development: Define technical career paths, performance benchmarks, and continuous learning opportunities in observability and SRE for mid-level and senior engineers.

Vendor Governance: Manage strategic relationships, licensing, contractual negotiations, and technical roadmaps with key observability and cloud partners (Splunk, Dynatrace, AWS, GitLab).

Required Qualifications & Experience Education & Overall Experience Bachelor’s or Master’s Degree in Computer Science, Software Engineering, Information Technology, or a related quantitative field.

15+ years of total experience in SRE, Systems/Platform Engineering, Infrastructure Architecture, and Enterprise Observability. Core Observability & Reliability Expertise Observability Mastery: Deep expertise architecting enterprise telemetry solutions using tools such as Splunk, Dynatrace, Datadog, Prometheus, Grafana, and OpenTelemetry.

AIOps & Automation: Proven track record implementing AIOps tools, automated incident remediation, and AI/ML-based anomaly detection engines.

SRE Leadership: Expert knowledge of SRE best practices, error budget policy enforcement, SLO modeling, and incident response frameworks. Technical & Systems Foundation Cloud & Modern Infrastructure: Strong hands-on architectural understanding of AWS services (EC2, ALB/NLB, Direct Connect, IAM, KMS, Transfer Family).

Legacy-to-Cloud Modernization: Practical experience executing "strangler fig" migration strategies off legacy tooling (Puppet, SVN, Xymon) to modern GitOps/IaC (GitLab CI/CD, Terraform, Ansible).

Data & Middleware: Experience monitoring and troubleshooting distributed middleware and database technologies (Apache Kafka, Cassandra) under heavy throughput.

FinTech & Secure Protocols: Understanding of secure file transfer (SFTP, AS2), mTLS, cryptographic key management (HSMs, Virtucrypt), and high-availability payment processing environments.

Preferred Certifications Observability: Splunk Certified Architect, Dynatrace Master / Professional, or equivalent.

Cloud & DevOps: AWS Certified Solutions Architect – Professional

SRE & Methodologies: Certified Site Reliability Engineer (SRE), ITIL v4 (Incident/Problem Management focus).

📌 Director – Observability, Response & Reliability (ORR) (Noida)
🏢 NetSpend
📍 Noida

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: director – observability, response & reliability (orr) (noida) / noida

Subscribe to this job alert:

Get the latest job offers by email for: director – observability, response & reliability (orr) (noida) / noida