06 Aug
|
Smbc Global Services
|
Chennai
06 Aug
Smbc Global Services
Chennai
Reporting to: Full Stack Engineer Lead
Department: Service Delivery Department (SDD)
Domain: Full Stack Engineering Shared Platform
Role Overview
As the DevOps & Production Engineering - GCC Lead based in our Chennai strategic technology center, you will be the key technical authority and leader responsible for ensuring the 24/7 availability, reliability, and top-tier performance of our next-generation digital banking web and mobile applications across the Asia Pacific region. Operating within a high-stakes, hybrid environment (Cloud and On-Premises), you will champion operational excellence and safeguard customer-facing banking journeys.
In this role, you will lead, scale, and mentor the Chennai team of DevOps and production support engineers, working closely with the counterparts in Singapore. This is a high-visibility, hands-on leadership role-you will drive critical incident resolution, orchestrate deep-dive root cause analyses alongside development teams, and transition operations from reactive troubleshooting to long-term proactive engineering solutions.
About the Opportunity
- Strategic Leadership: Own and scale the 24/7 global operational support strategy for a premier digital banking platform from our Chennai technology hub.
- Hands-on Triage: Actively lead high-severity incident troubleshooting sessions, guiding teams to deliver immediate technical workarounds and permanent fixes.
- Enterprise Scale Architecture: Oversee complex multi-tier ecosystems featuring web frontends, mobile applications (iOS/Android), extensive microservices, and multiple relational/NoSQL databases across hybrid cloud infrastructures.
- Cross-Regional Collaboration: Manage and unify diverse, multi-regional onshore and offshore support engineering teams, coordinating seamless follow-the-sun handoffs between Chennai, Singapore, and other tech hubs.
- SRE & Automation Transformation: Drive the reduction of operational toil by embedding modern Site Reliability Engineering (SRE) and automation principles into traditional support models.
Key Responsibilities
Incident Management & Hands-On Troubleshooting
- Crisis Command: Actively lead technical troubleshooting sessions for critical (Severity 1 and 2) production incidents affecting web, mobile, backend services, and critical data layers.
- Rapid Resolution & Reporting: Ensure rapid restoration of services while maintaining clear, real-time executive communication and business stakeholder updates during major incidents.
- Root Cause Elimination: Partner closely with Development, Cloud Infrastructure, and
- DevOps teams to conduct detailed Post-Mortems and Blameless Root Cause Analyses (RCA), tracking temporary workarounds through to permanent software or architectural bug fixes.
Team Leadership & Global Operations
- Follow-the-Sun Governance: Manage and optimize multi-geographical support squads operating across Chennai and Singapore to guarantee seamless, sustainable 24/7 production coverage.
- Capability Building: Mentor and elevate the technical capabilities of support and DevOps engineers, establishing clear engineering career paths and runbook proficiencies.
- Operational Readiness: Define, track, and regularly report on critical SLAs, OLAs, and service metrics including MTTR (Mean Time to Resolution) and MTTD (Mean Time to Detection) to senior IT leadership.
Hybrid Platform & Database Reliability
- Hybrid Infrastructure Support: Manage and troubleshoot applications distributed across both traditional On-Premises enterprise datacenters and Cloud native environments (Azure/AWS).
- Multi-Database Governance: Oversee operational health, query performance, and failover/replication mechanisms across multiple database backends (SQL Server, Oracle, PostgreSQL, NoSQL).
- Observability Engineering: Drive the design and enhancement of enterprise monitoring, tracing, and logging solutions (e.g., Dynatrace, Datadog, Splunk, ELK) to systematically detect anomalies before they impact end-users.
Process Optimization & Automation
- Toil Elimination: Identify repetitive manual tasks and champion an automation-first approach, developing scripts to automate standard health-checks, recovery procedures, and deployments.
- ITIL Excellence: Govern the implementation of high-standard ITIL frameworks covering Incident, Problem, Change, and Release management tailored for rapid-deployment environments.
- Disaster Recovery Strategy: Plan, lead, and execute complex regular Disaster Recovery (DR) and business continuity drills for critical banking channels.
Security, Risk & Compliance
- Access Control: Enforce strict access management, principle of least privilege, and secure credential handling (e.g., HashiCorp Vault) across all production tiers.
- Regulatory Adherence: Ensure all production operations, handling of data, and incident responses strictly comply with regional banking security regulations
- Audit Readiness: Represent production operations in security audits, leading immediate remediation initiatives for identified technical or procedural vulnerabilities.
Required Qualifications
Technical Expertise
Must demonstrate high proficiency in at least 4 of the following areas:
- Enterprise-Scale Application Architecture: Deep conceptual understanding of multitier web/mobile applications, distributed microservices architectures, RESTful APIs, and enterprise API Gateways to quickly isolate components during an incident.
- Hybrid-Cloud Infrastructure Exposure:
Extensive exposure to managing and supporting enterprise-scale banking operations running across Azure (preferred) or AWS, alongside traditional On-Premises corporate data centers.
- SQL & Shell Scripting Proficiency: Advanced hands-on skills in writing SQL queries across multiple enterprise databases (Oracle, SQL Server, PostgreSQL) and high proficiency in Shell Scripting (Bash/Unix) to manipulate logs, query data anomalies, and run terminallevel triage.
- Programming Language Knowledge: Strong code-reading and diagnostic knowledge of Python or Java to analyze error stack traces and partner with developers on fixes (no active coding or application development required).
- Containerization & Orchestration Exposure: Operational exposure to containerized applications running on Docker and Kubernetes (AKS/EKS) to navigate clusters, check pod states, and extract logs during incidents.
- Monitoring & Observability Platforms: Expert usage of enterprise observability platforms (Datadog, Splunk, Dynatrace, Prometheus, Grafana) for multi-tier log correlation, telemetry metrics analysis, and proactive performance bottleneck detection.
- CI/CD & DevOps Workflow Familiarity: Practical knowledge of DevOps pipelines (GitLab CI, Azure DevOps, or Jenkins) and modern AI engineering assistants to quickly trace and isolate deployment-related regressions (no active pipeline development required).
Experience & Education
- 14-17 years of overall experience within Production Support, Application Support, Systems Engineering, or Site Reliability Engineering (SRE) roles.
- 4+ years in an engineering leadership or people management role managing multiregional teams under a 24/7 coverage model.
- Enterprise Financial Background: Extensive experience supporting critical core or digital banking systems inside a tier-1 financial institution or highly regulated fintech environment.
- Education: Bachelors degree in Computer Science, Information Technology, or a related quantitative technical field (or equivalent practical experience).
Professional Qualities
- Crisis Composure: Calm under intense pressure; able to logically dissect complex system failures while under tight recovery timelines.
- Strategic & Analytical Mindset: Looks beyond the immediate fix to systematically evaluate the architectural flaws causing chronic operational pain.
- Global Leadership Polish: Solid emotional intelligence; adept at motivating and driving engineering squads across diverse cultural and geographic boundaries (Chennai, Singapore, APAC) while setting clear, uncompromising performance expectations.
- Executive Communicator: Capable of translating highly technical, complex incident data into concise, business-impact summaries for regional C-suite and Managing Director review.
- Automation-First Philosophy: Deeply intolerant of repetitive manual efforts ("toil"); consistently steers the engineering team toward programmatic, self-healing remedies.
📌 DevOps & Production Lead (Chennai)
🏢 Smbc Global Services
📍 Chennai