DevOps & Production Lead (Chennai)

DevOps & Production Lead (Chennai)

06 Aug
|
Smbc Global Services
|
Chennai

06 Aug

Smbc Global Services

Chennai

Reporting to: Full Stack Engineer Lead

Department: Service Delivery Department (SDD)

Domain: Full Stack Engineering Shared Platform

Role Overview

As the DevOps & Production Engineering - GCC Lead based in our Chennai strategic technology center, you will be the key technical authority and leader responsible for ensuring the 24/7 availability, reliability, and top-tier performance of our next-generation digital banking web and mobile applications across the Asia Pacific region. Operating within a high-stakes, hybrid environment (Cloud and On-Premises), you will champion operational excellence and safeguard customer-facing banking journeys.

In this role, you will lead, scale, and mentor the Chennai team of DevOps and production support engineers, working closely with the counterparts in Singapore. This is a high-visibility, hands-on leadership role-you will drive critical incident resolution, orchestrate deep-dive root cause analyses alongside development teams, and transition operations from reactive troubleshooting to long-term proactive engineering solutions.

About the Opportunity

- Strategic Leadership: Own and scale the 24/7 global operational support strategy for a premier digital banking platform from our Chennai technology hub.
- Hands-on Triage: Actively lead high-severity incident troubleshooting sessions, guiding teams to deliver immediate technical workarounds and permanent fixes.
- Enterprise Scale Architecture: Oversee complex multi-tier ecosystems featuring web frontends, mobile applications (iOS/Android), extensive microservices, and multiple relational/NoSQL databases across hybrid cloud infrastructures.
- Cross-Regional Collaboration: Manage and unify diverse, multi-regional onshore and offshore support engineering teams, coordinating seamless follow-the-sun handoffs between Chennai, Singapore, and other tech hubs.
- SRE & Automation Transformation: Drive the reduction of operational toil by embedding modern Site Reliability Engineering (SRE) and automation principles into traditional support models.

Key Responsibilities

Incident Management & Hands-On Troubleshooting

- Crisis Command: Actively lead technical troubleshooting sessions for critical (Severity 1 and 2) production incidents affecting web, mobile, backend services, and critical data layers.
- Rapid Resolution & Reporting: Ensure rapid restoration of services while maintaining clear, real-time executive communication and business stakeholder updates during major incidents.
- Root Cause Elimination: Partner closely with Development, Cloud Infrastructure, and
- DevOps teams to conduct detailed Post-Mortems and Blameless Root Cause Analyses (RCA), tracking temporary workarounds through to permanent software or architectural bug fixes.

Team Leadership & Global Operations

- Follow-the-Sun Governance: Manage and optimize multi-geographical support squads operating across Chennai and Singapore to guarantee seamless, sustainable 24/7 production coverage.




- Capability Building: Mentor and elevate the technical capabilities of support and DevOps engineers, establishing clear engineering career paths and runbook proficiencies.
- Operational Readiness: Define, track, and regularly report on critical SLAs, OLAs, and service metrics including MTTR (Mean Time to Resolution) and MTTD (Mean Time to Detection) to senior IT leadership.

Hybrid Platform & Database Reliability

- Hybrid Infrastructure Support: Manage and troubleshoot applications distributed across both traditional On-Premises enterprise datacenters and Cloud native environments (Azure/AWS).
- Multi-Database Governance: Oversee operational health, query performance, and failover/replication mechanisms across multiple database backends (SQL Server, Oracle, PostgreSQL, NoSQL).
- Observability Engineering: Drive the design and enhancement of enterprise monitoring, tracing, and logging solutions (e.g., Dynatrace, Datadog, Splunk, ELK) to systematically detect anomalies before they impact end-users.

Process Optimization & Automation

- Toil Elimination: Identify repetitive manual tasks and champion an automation-first approach, developing scripts to automate standard health-checks, recovery procedures, and deployments.
- ITIL Excellence: Govern the implementation of high-standard ITIL frameworks covering Incident, Problem, Change, and Release management tailored for rapid-deployment environments.
- Disaster Recovery Strategy: Plan, lead, and execute complex regular Disaster Recovery (DR) and business continuity drills for critical banking channels.

Security, Risk & Compliance

- Access Control: Enforce strict access management, principle of least privilege, and secure credential handling (e.g., HashiCorp Vault) across all production tiers.
- Regulatory Adherence: Ensure all production operations, handling of data, and incident responses strictly comply with regional banking security regulations
- Audit Readiness: Represent production operations in security audits, leading immediate remediation initiatives for identified technical or procedural vulnerabilities.

Required Qualifications

Technical Expertise

Must demonstrate high proficiency in at least 4 of the following areas:

- Enterprise-Scale Application Architecture: Deep conceptual understanding of multitier web/mobile applications, distributed microservices architectures, RESTful APIs, and enterprise API Gateways to quickly isolate components during an incident.
- Hybrid-Cloud Infrastructure Exposure:



Extensive exposure to managing and supporting enterprise-scale banking operations running across Azure (preferred) or AWS, alongside traditional On-Premises corporate data centers.
- SQL & Shell Scripting Proficiency: Advanced hands-on skills in writing SQL queries across multiple enterprise databases (Oracle, SQL Server, PostgreSQL) and high proficiency in Shell Scripting (Bash/Unix) to manipulate logs, query data anomalies, and run terminallevel triage.
- Programming Language Knowledge: Strong code-reading and diagnostic knowledge of Python or Java to analyze error stack traces and partner with developers on fixes (no active coding or application development required).
- Containerization & Orchestration Exposure: Operational exposure to containerized applications running on Docker and Kubernetes (AKS/EKS) to navigate clusters, check pod states, and extract logs during incidents.
- Monitoring & Observability Platforms: Expert usage of enterprise observability platforms (Datadog, Splunk, Dynatrace, Prometheus, Grafana) for multi-tier log correlation, telemetry metrics analysis, and proactive performance bottleneck detection.
- CI/CD & DevOps Workflow Familiarity: Practical knowledge of DevOps pipelines (GitLab CI, Azure DevOps, or Jenkins) and modern AI engineering assistants to quickly trace and isolate deployment-related regressions (no active pipeline development required).

Experience & Education

- 14-17 years of overall experience within Production Support, Application Support, Systems Engineering, or Site Reliability Engineering (SRE) roles.
- 4+ years in an engineering leadership or people management role managing multiregional teams under a 24/7 coverage model.
- Enterprise Financial Background: Extensive experience supporting critical core or digital banking systems inside a tier-1 financial institution or highly regulated fintech environment.
- Education: Bachelors degree in Computer Science, Information Technology, or a related quantitative technical field (or equivalent practical experience).

Professional Qualities

- Crisis Composure: Calm under intense pressure; able to logically dissect complex system failures while under tight recovery timelines.
- Strategic & Analytical Mindset: Looks beyond the immediate fix to systematically evaluate the architectural flaws causing chronic operational pain.
- Global Leadership Polish: Solid emotional intelligence; adept at motivating and driving engineering squads across diverse cultural and geographic boundaries (Chennai, Singapore, APAC) while setting clear, uncompromising performance expectations.
- Executive Communicator: Capable of translating highly technical, complex incident data into concise, business-impact summaries for regional C-suite and Managing Director review.
- Automation-First Philosophy: Deeply intolerant of repetitive manual efforts ("toil"); consistently steers the engineering team toward programmatic, self-healing remedies.

📌 DevOps & Production Lead (Chennai)
🏢 Smbc Global Services
📍 Chennai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: devops & production lead (chennai) / chennai

Subscribe to this job alert:

Get the latest job offers by email for: devops & production lead (chennai) / chennai