02 Oct
|
ALTIMETRIK
|
Bengaluru
02 Oct
ALTIMETRIK
Bengaluru
Greetings from Altimetrik !!!
Hiring Application support & SRE Lead Professionals for our concern
Experience : 8 years to 15 years
Work location: Bangalore ( Weekly thirce in office)
Notice Period : Less than 20 days (Serving Notice) /Immediate
Interview Mode : Microsoft Teams
:
Application Support Lead
Role Purpose
Own the day-to-day health of the production support function for a high-volume card payments platform. The Lead is accountable for service quality end to end incident outcomes, problem closure, shift coverage and the capability of the L2/L3 engineers and is the senior support voice in major incidents, change reviews and client-facing escalations.
Responsibilities
- Own overall service quality for application support: SLA performance, P1 restoration targets, backlog health and escalation discipline across the L2 and L3 tiers
- Act as incident commander for P1 and major incidents coordinate responders, control communications to clients and internal stakeholders, and decide when to escalate, invoke DR, or hold for a controlled fix
- Own problem management: ensure RCAs are produced to a standard auditors and clients can read, and drive preventative actions to closure rather than repeat firefighting
- Own the 247365 roster shift design, handover quality, on-call fairness and coverage across UK/India/APAC time zones
- Line-manage and develop L2/L3 engineers: objectives, one-to-ones, performance reviews, coaching, and a transparent skills path from L2 to L3 to SME
- Run recruitment and onboarding for the team; build ramp plans that get new joiners productive on the payments domain quickly
- Represent support at CAB and release governance — assess change risk, approve pre-approved vs. exception routes, and sign off post-deployment validation
- Own the monitoring and alerting strategy: reduce noise, close detection gaps, and set thresholds so issues are caught before clients report them
- Own the knowledge estate — runbooks, KEDB, SOPs and knowledge articles — and enforce the standard that no process sits with one person
- Report on service health to senior management and clients: incident trends, SLA attainment, top recurring issues, and the plan to remove them
- Manage escalations from client account teams, scheme-facing teams and tier 1 clients, translating technical cause into an accurate, reassuring answer
- Own shift-left: move resolvable issues down from L3 to L2 and from L2 to self-service, backed by automation rather than extra headcount
- Coordinate with Engineering, DBAs, Infosec, Compliance and third-party vendors on cross-team incidents, patching cycles and platform changes
- Maintain audit and regulatory readiness for the support function — evidence trails, access control, change records, and operational resilience obligations
Reliability Engineering (SRE) :
- Define and own SLIs, SLOs and error budgets for the platform’s critical journeys — authorisation, clearing, settlement, file delivery and reporting — and use error budget burn to arbitrate between change velocity and stability
- Drive systematic toil reduction: measure the manual, repetitive effort inside the support queue and set a target for how much of it is automated away each quarter
- Build and own automation for recurring operational work — runbook automation, self-healing jobs, diagnostic scripts and auto-remediation for known errors
- Own observability as a product: golden-signal dashboards, structured logging standards, trace coverage, synthetic monitoring of key transaction paths, and alerts tied to symptoms customers feel rather than raw resource metrics
- Run blameless post-incident reviews; track MTTD/MTTA/MTTR and incident recurrence as managed metrics, and feed reliability debt into the engineering backlog with evidence
- Own capacity and performance management — throughput headroom for peak card volumes, queue depth trends, database growth, and early warning before saturation
- Plan and execute resilience testing: DR invocation and failover drills, backup restore validation, and controlled game days against production-like environments
- Set and enforce production readiness criteria for new services and releases — monitoring, runbook, rollback path and support model in place before go-live
- Partner with Engineering on reliability by design: input to architecture reviews, retry and timeout behaviour, graceful degradation, and dependency failure modes
- Use infrastructure-as-code and CI/CD to make operational changes repeatable, reviewable and auditable under CAB
Technical Skills
- Deep incident and problem management practice in a 247 regulated production environment; ITIL-aligned incident, problem, change and major incident handling
- Hands-on credibility across the L3 stack — able to review and challenge diagnostics rather than only receive them
- SQL (MySQL, PostgreSQL, SQL Server) — review and approve scripts and stored procedures destined for production
- PowerShell and scripting for automation; .NET log and code-level familiarity sufficient to direct an investigation
- Observability tooling — Coralogix, CloudWatch, Grafana, PagerDuty — including alert design, dashboard ownership and noise reduction
- AWS (ECS, EKS, RDS, API Gateway) and an understanding of environment topology, load-balanced services and failover behaviour
- Payments domain — card issuing and processing lifecycle, Visa/Mastercard message flows, ISO 8583 / ISO 20022, clearing and settlement, chargebacks and disputes, tokenisation (VTS, MDES)
- File and data pipelines — XML/CSV reporting, SFTP transfers, schema and encoding failure modes
- PCI DSS controls and FCA operational resilience — able to own the support function’s obligations, not just work within them
- Jira Service Management and Confluence — service design, queue and workflow configuration, reporting
- SRE practice — SLI/SLO definition, error budget policy, toil measurement, blameless postmortems, production readiness reviews
- Automation tooling — infrastructure-as-code (Terraform or CloudFormation), CI/CD pipelines, and scripted remediation
- Capacity, performance and resilience testing, including DR failover and backup restore validation
Good to Have: Kubernetes and container platform operations at scale • Distributed tracing (OpenTelemetry) • Chaos engineering / game day facilitation • AI-assisted troubleshooting tooling in a controlled environment • Vendor and third-party management • Scheme compliance exposure (QMR/GOC, interchange mandates) • Payments/fintech background
📌 Application Support Lead (Bengaluru)
🏢 ALTIMETRIK
📍 Bengaluru