09 Sep
|
Kotak Mahindra Bank
|
Hyderabad
09 Sep
Kotak Mahindra Bank
Hyderabad
Tech Ops Engineering II-SUPPORT SERVICES-CTO - In House Engineering Job Title: Production Site Reliability Engineer (SRE) – Digital Payments Role Overview We are seeking a highly technical and driven Production SRE Engineer to manage and monitor mission-critical payment platforms including UPI, IMPS, and Payment Hub systems & various Payment Applications . The role focuses on ensuring high availability, low latency, and seamless transaction experience for customers. The incumbent will collaborate with cross-functional teams (Engineering, Business, Compliance) and external regulators ( RBI, NPCI ) to maintain resilient and scalable payment infrastructure.
Key Responsibilities Production Support &
- Incident Management Provide L2/L3 production support for UPI, IMPS, and Payment Hub platforms & various Payment Applications . Diagnose, triage, and resolve transaction failures, timeouts, and API disruptions . Lead and participate in Major Incident Management (MIM) calls and ensure timely stakeholder communication. Manage incidents, service requests, and problem tickets via Jira, ServiceNow . Provide regular updates to internal stakeholders and regulatory bodies (NPCI/RBI) during critical issues.
Reliability
Engineering &
- RCA Perform deep-dive Root Cause Analysis (RCA) for recurring payment and system issues. Implement preventive and corrective measures to improve system stability. Drive SRE best practices including error budgets, SLIs/SLOs, and system resilience.
Experience in managing DR Drills &
- Documentations. Reviewing the SOPs & its relative documentations. Monitoring, Observability &
- System Engineering Monitor key performance indicators: Transaction success rates Latency and response times Failure trends and retries Build and maintain dashboards using: ELK Stack,
Grafana, Kibana, Splunk, Datadog, Prometheus Establish proactive alerting and anomaly detection mechanisms. Work closely with engineering teams to design and optimize: High-throughput payment switches Routing logic Settlement and reconciliation systems Understand and support UPI architecture, IMPS rails, and payment orchestration layers & various Payment Applications . Trace end-to-end transaction lifecycle across distributed systems.
External
Partner &
- Regulatory Coordination Coordinate with NPCI, partner banks, and TPAPs during outages, reconciliation issues, or network disruptions. Lead integrations and ensure seamless onboarding of ecosystem participants.
Ensure compliance with: RBI guidelines and data localization mandates NPCI operational and technical standards Technical Skills &
- Expertise Payments Domain Knowledge Solid expertise in: UPI architecture and flows IMPS rails Payment gateway / switch systems Payment Hub orchestration & various Payment Applications.
Core Technical Skills
Advanced SQL proficiency (joins, aggregations, stored procedures) Strong hands-on experience in: Linux/UNIX systems administration Shell scripting Ability to: Read , Write and interpret All types documentation (SOPs, workflows, etc.) Understand database schemas Analyse system architecture and latency Monitoring &
- Observability Tools Hands-on expertise with:
ELK Stack (Elasticsearch, Logstash, Kibana) Grafana, Prometheus Splunk, Datadog DevOps &
- Cloud Experience with: CI/CD pipelines, Containerization (Docker, Kubernetes) Cloud platforms: AWS / GCP / Azure Key Competencies Strong problem-solving and analytical skills High ownership in production environments Ability to work under pressure in real-time systems Strong stakeholder communication and coordination Focus on reliability, scalability, and performance Must have can do, takes initiative, Drives end to end deliverables. Proactive, solution-oriented mindset with ownership to resolve production issues under pressure. Ability to clearly articulate incidents, updates, and RCA to stakeholders, leadership, and regulators. Works effectively with cross-functional teams (engineering, product, partners, regulators). Structured thinking to diagnose complex system failures and drive long-term fixes. Ability to stay calm and effective during high-severity incidents and critical outages. Knowledge of PCI-DSS compliance, Financial data governance & security best practices Quickly adapts to changing technologies, incidents, and regulatory requirements in a fast-evolving payments ecosystem. Precision in analysing logs, transactions, and system behaviour to avoid critical errors in production. Effectively manage multiple incidents, tasks, and escalations in a high-pressure environment. Ability to handle expectations and coordinate with internal teams, partners, and regulators efficiently. Takes ownership to make quick, informed decisions during outages or critical production incidents & communications to various Stake holders including Regulatory.
Experience Level Executive Level
📌 Production Site Reliability Engineer - Digital Payments (Hyderabad)
🏢 Kotak Mahindra Bank
📍 Hyderabad