28 Sep
|
ALTIMETRIK
|
Bengaluru
28 Sep
ALTIMETRIK
Bengaluru
Application Support Lead
Role Purpose
Own the day-to-day health of the production support function for a high-volume card payments platform. The Lead is accountable for service quality end to end incident outcomes, problem closure, shift coverage and the capability of the L2/L3 engineers — and is the senior support voice in major incidents, change reviews and client-facing escalations.
Responsibilities
•Own overall service quality for application support: SLA performance, P1 restoration targets, backlog health and escalation discipline across the L2 and L3 tiers
•Act as incident commander for P1 and major incidents — coordinate responders, control communications to clients and internal stakeholders, and decide when to escalate, invoke DR, or hold for a controlled fix
•Own problem management: ensure RCAs are produced to a standard auditors and clients can read, and drive preventative actions to closure rather than repeat firefighting
•Own the 247365 roster — shift design, handover quality, on-call fairness and coverage across UK/India/APAC time zones
•Line-manage and develop L2/L3 engineers: objectives, one-to-ones, performance reviews, coaching, and a clear skills path from L2 to L3 to SME
•Run recruitment and onboarding for the team; build ramp plans that get new joiners productive on the payments domain quickly
•Represent support at CAB and release governance — assess change risk, approve pre-approved vs. exception routes, and sign off post-deployment validation
•Own the monitoring and alerting strategy: reduce noise, close detection gaps, and set thresholds so issues are caught before clients report them
•Own the knowledge estate — runbooks, KEDB, SOPs and knowledge articles — and enforce the standard that no process sits with one person
•Report on service health to senior management and clients: incident trends, SLA attainment, top recurring issues, and the plan to remove them
•Manage escalations from client account teams, scheme-facing teams and tier 1 clients, translating technical cause into an accurate, reassuring answer
•Own shift-left: move resolvable issues down from L3 to L2 and from L2 to self-service, backed by automation rather than extra headcount
•Coordinate with Engineering, DBAs, Infosec,
Compliance and third-party vendors on cross-team incidents, patching cycles and platform changes
•Maintain audit and regulatory readiness for the support function — evidence trails, access control, change records, and operational resilience obligations
Reliability Engineering (SRE)
•Define and own SLIs, SLOs and error budgets for the platform’s critical journeys — authorisation, clearing, settlement, file delivery and reporting — and use error budget burn to arbitrate between change velocity and stability
•Drive systematic toil reduction: measure the manual, repetitive effort inside the support queue and set a target for how much of it is automated away each quarter
•Build and own automation for recurring operational work — runbook automation, self-healing jobs, diagnostic scripts and auto-remediation for known errors
•Own observability as a product: golden-signal dashboards, structured logging standards, trace coverage, synthetic monitoring of key transaction paths, and alerts tied to symptoms customers feel rather than raw resource metrics
•Run blameless post-incident reviews; track MTTD/MTTA/MTTR and incident recurrence as managed metrics, and feed reliability debt into the engineering backlog with evidence
•Own capacity and performance management — throughput headroom for peak card volumes, queue depth trends, database growth, and early warning before saturation
•Plan and execute resilience testing: DR invocation and failover drills, backup restore validation, and controlled game days against production-like environments
•Set and enforce production readiness criteria for new services and releases — monitoring, runbook, rollback path and support model in place before go-live
•Partner with Engineering on reliability by design: input to architecture reviews, retry and timeout behaviour, graceful degradation,
and dependency failure modes
•Use infrastructure-as-code and CI/CD to make operational changes repeatable, reviewable and auditable under CAB
Technical Skills
•Deep incident and problem management practice in a 247 regulated production environment; ITIL-aligned incident, problem, change and major incident handling
•Hands-on credibility across the L3 stack — able to review and challenge diagnostics rather than only receive them
•SQL (MySQL, PostgreSQL, SQL Server) — review and approve scripts and stored procedures destined for production
•PowerShell and scripting for automation; .NET log and code-level familiarity sufficient to direct an investigation
•Observability tooling — Coralogix, CloudWatch, Grafana, PagerDuty — including alert design, dashboard ownership and noise reduction
•AWS (ECS, EKS, RDS, API Gateway) and an understanding of environment topology, load-balanced services and failover behaviour
•Payments domain — card issuing and processing lifecycle, Visa/Mastercard message flows, ISO 8583 / ISO 20022, clearing and settlement, chargebacks and disputes, tokenisation (VTS, MDES)
•File and data pipelines — XML/CSV reporting, SFTP transfers, schema and encoding failure modes
•PCI DSS controls and FCA operational resilience — able to own the support function’s obligations, not just work within them
•Jira Service Management and Confluence — service design, queue and workflow configuration, reporting
•SRE practice — SLI/SLO definition, error budget policy, toil measurement, blameless postmortems, production readiness reviews
•Automation tooling — infrastructure-as-code (Terraform or CloudFormation), CI/CD pipelines, and scripted remediation
•Capacity, performance and resilience testing, including DR failover and backup restore validation
Good to Have: Kubernetes and container platform operations at scale • Distributed tracing (OpenTelemetry) • Chaos engineering / game day facilitation • AI-assisted troubleshooting tooling in a controlled setting • Vendor and third-party management • Scheme compliance exposure (QMR/GOC, interchange mandates) • Payments/fintech background
📌 Application Support Lead (Bengaluru)
🏢 ALTIMETRIK
📍 Bengaluru