11 Sep
|
HDFC Bank
|
Bengaluru
11 Sep
HDFC Bank
Bengaluru
Job Title:
Lead Site Reliability Engineer
Job Details:
Business Unit: Tech & Digital
Team: Ent Factory-Channels, Mobility, Payments
Reports to: SRE Manager
Location: Mumbai, Chennai, Gurgaon & Bangalore
Role Type: Non-Supervisory
No of direct reportees: 0
Travel Required: No
Job Band Range: D1/D2
JD Created date: 28th Jan 2023
Job Purpose:
Analyzing, troubleshooting, and designing vital services, platforms, and infrastructure on GCP with a focus on reliability, scalability, resilience, security, and performance.
Lead and Mentor a team of SRE engineers.
Job Responsibilities:
Execute reliability initiatives for the team and organization.
Mentor and lead a team of SRE engineers.
Help build a Site Reliability Engineering culture by sharing best practices, approaches, documentation, and code with other engineering teams.
Solid understanding of observability tools and ability to express reliability metrics via observability.
Define KPI in the form of RPO/RTO/SLI/SLO/Error Budget.
Apply automation and software to any manually performed tasks or system parts.
Troubleshoot complicated, cross-platform issues handling OS, Networking, Database in a cloud-based SaaS workplace and manage live production incidents.
Monitor application performance, take steps to improve overall application performance and stability, and follow through with implementation.
Conduct system analysis, configuration management, and develop improvements for system software performance, availability, and reliability.
Design, write, ship, and motivate the creation of software and systems to increase observability, product reliability, and organizational efficiency.
Maintain and monitor deployment, orchestration of servers, docker containers, databases, and general backend infrastructure.
Develop Run Books/Standard Operating Procedure for recurring Production issues and work on permanent solutions.
Perform Incident Analysis regularly to prevent and find long-term solutions for Incidents.
Educational Qualifications:
B Tech in Computer Science or related discipline preferred.
Key Skills:
Experience in monitoring and analyzing infrastructure performance using standard performance monitoring tools.
Demonstrable experience in Containerization (Docker) and orchestration (Kubernetes).
Experience with Infrastructure As Code (Terraform, Cloud Formation, Ansible).
Knowledge and proven hands-on experience in large-scale databases and distributed technologies, such as Kafka and Confluent Platform Kafka.
Basic programming and scripting skills.
Solid understanding of at least 2 observability technologies.
Experience Required:
Total Yrs of experience: 11-13
Major Stakeholders:
Internal: Product Manager from Digital Factory, Business Analyst from BTG team, Incident Management team, Development Team.
Skills
Docker,Kubernates,AWS,Monitoring,terraform,ansible,networking,reliability engineering
About HDFC Bank
HDFC Bank is one of India’s leading private banks. The Housing Development Finance Corporation Limited or HDFC Ltd was among the first financial institutions in India to receive an “in principle” approval from the Reserve Bank of India (RBI) to set up a bank in the private sector. This was done as part of RBI’s policy for liberalization of the Indian banking industry in 1994. HDFC Bank was incorporated in August 1994 in the name of HDFC Bank Limited, with its registered office in Mumbai, India. The bank commenced operations as a Scheduled Commercial Bank in January 1995. As of March 31, 2026, the Bank’s distribution network was at 9,689 branches and 21,172 ATMs across 4,175 cities / towns as against 9,455 branches and 21,139 ATMs across 4,150 cities / towns as of March 31, 2025. 50% of the branches are in semi-urban and rural areas. In addition, the Bank has 14,400 business correspondents, which are primarily manned by Common Service Centres (CSC). The Bank’s international operations comprise five branch
View more
📌 Tech & Digital-Lead Site Reliability Engineer (Bengaluru)
🏢 HDFC Bank
📍 Bengaluru