02 Oct
|
ALTIMETRIK
|
Bengaluru
02 Oct
ALTIMETRIK
Bengaluru
Greetings from Altimetrik !!!
Hiring Application support & SRE Lead Professionals for our concern
Experience : 8 years to 15 years
Work location: Bangalore ( Weekly thirce in office)
Notice Period : Less than 20 days (Serving Notice) /Immediate
Interview Mode : Microsoft Teams
Job Description :
Application Support Lead
Role Purpose
Own the day-to-day health of the production support function for a high-volume card payments platform. The Lead is accountable for service quality end to end incident outcomes, problem closure, shift coverage and the capability of the L2/L3 engineers and is the senior support voice in major incidents, change reviews and client-facing escalations.
Responsibilities
• Own overall service quality for application support: SLA performance, P1 restoration targets, backlog health and escalation discipline across the L2 and L3 tiers
• Act as incident commander for P1 and major incidents coordinate responders, control communications to clients and internal stakeholders, and decide when to escalate, invoke DR, or hold for a controlled fix
• Own problem management: ensure RCAs are produced to a standard auditors and clients can read, and drive preventative actions to closure rather than repeat firefighting
• Own the 247365 roster shift design, handover quality, on-call fairness and coverage across UK/India/APAC time zones
• Line-manage and develop L2/L3 engineers: objectives, one-to-ones, performance reviews, coaching, and a clear skills path from L2 to L3 to SME
• Run recruitment and onboarding for the team; build ramp plans that get current joiners productive on the payments domain quickly
• Represent support at CAB and release governance — assess change risk, approve pre-approved vs. exception routes, and sign off post-deployment validation
• Own the monitoring and alerting strategy: reduce noise, close detection gaps, and set thresholds so issues are caught before clients report them
• Own the knowledge estate — runbooks, KEDB, SOPs and knowledge articles — and enforce the standard that no process sits with one person
• Report on service health to senior management and clients: incident trends, SLA attainment, top recurring issues, and the plan to remove them
• Manage escalations from client account teams, scheme-facing teams and tier 1 clients, translating technical cause into an accurate,
reassuring answer
• Own shift-left: move resolvable issues down from L3 to L2 and from L2 to self-service, backed by automation rather than extra headcount
• Coordinate with Engineering, DBAs, Infosec, Compliance and third-party vendors on cross-team incidents, patching cycles and platform changes
• Maintain audit and regulatory readiness for the support function — evidence trails, access control, change records, and operational resilience obligations
Reliability Engineering (SRE) :
• Define and own SLIs, SLOs and error budgets for the platform’s critical journeys — authorisation, clearing, settlement, file delivery and reporting — and use error budget burn to arbitrate between change velocity and stability
• Drive systematic toil reduction: measure the manual, repetitive effort inside the support queue and set a target for how much of it is automated away each quarter
• Build and own automation for recurring operational work — runbook automation, self-healing jobs, diagnostic scripts and auto-remediation for known errors
• Own observability as a product: golden-signal dashboards, structured logging standards, trace coverage, synthetic monitoring of key transaction paths, and alerts tied to symptoms customers feel rather than raw resource metrics
• Run blameless post-incident reviews; track MTTD/MTTA/MTTR and incident recurrence as managed metrics, and feed reliability debt into the engineering backlog with evidence
• Own capacity and performance management — throughput headroom for peak card volumes, queue depth trends, database growth, and early warning before saturation
• Plan and execute resilience testing: DR invocation and failover drills, backup restore validation, and controlled game days against production-like environments
• Set and enforce production readiness criteria for new services and releases — monitoring, runbook, rollback path and support model in place before go-live
• Partner with Engineering on reliability by design: input to architecture reviews, retry and timeout behaviour, graceful degradation, and dependency failure modes
• Use infrastructure-as-code and CI/CD to make operational changes repeatable, reviewable and auditable under CAB
Technical Skills
• Deep incident and problem management practice in a 247 regulated production environment; ITIL-aligned incident, problem, change and major incident handling
• Hands-on credibility across the L3 stack — able to review and challenge diagnostics rather than only receive them
• SQL (MySQL, PostgreSQL, SQL Server) — review and approve scripts and stored procedures destined for production
• PowerShell and scripting for automation; .NET log and code-level familiarity sufficient to direct an investigation
• Observability tooling — Coralogix, CloudWatch, Grafana, PagerDuty — including alert design, dashboard ownership and noise reduction
• AWS (ECS, EKS, RDS, API Gateway) and an understanding of environment topology, load-balanced services and failover behaviour
• Payments domain — card issuing and processing lifecycle, Visa/Mastercard message flows, ISO 8583 / ISO 20022, clearing and settlement, chargebacks and disputes, tokenisation (VTS, MDES)
• File and data pipelines — XML/CSV reporting, SFTP transfers, schema and encoding failure modes
• PCI DSS controls and FCA operational resilience — able to own the support function’s obligations, not just work within them
• Jira Service Management and Confluence — service design, queue and workflow configuration, reporting
• SRE practice — SLI/SLO definition, error budget policy, toil measurement, blameless postmortems, production readiness reviews
• Automation tooling — infrastructure-as-code (Terraform or CloudFormation), CI/CD pipelines, and scripted remediation
• Capacity, performance and resilience testing, including DR failover and backup restore validation
Good to Have: Kubernetes and container platform operations at scale • Distributed tracing (OpenTelemetry) • Chaos engineering / game day facilitation • AI-assisted troubleshooting tooling in a controlled environment • Vendor and third-party management • Scheme compliance exposure (QMR/GOC, interchange mandates) • Payments/fintech background
📌 Application Support Lead (Bengaluru)
🏢 ALTIMETRIK
📍 Bengaluru