12 Sep
|
HDFC Bank
|
Bengaluru
12 Sep
HDFC Bank
Bengaluru
Job Title:
Senior Site Reliability Engineer
Job Details:
Business Unit: Tech & Digital
Team: DTIT - Enterprise Factory
Reports to: Lead Site Reliability Engineer
Location: Mumbai, Chennai, Gurgaon & Bangalore
Role Type: Individual Contributor
No of direct reportees: NIL
Travel Required: No
Job Band Range: E4
Job Purpose:
Analysing, troubleshooting, and designing vital services, platforms, and infrastructure while always thinking about reliability, scalability, resilience, security, and performance.
Job Responsibili****ties:
Help build a Site Reliability Engineering culture by sharing best practices, approaches, documentation, and code with other engineering teams.
Apply automation and software to any tasks or parts of the system which are performed manually.
Able to troubleshoot complicated, cross-platform issues handling OS, Networking, Database in a cloud-based SaaS environment and handle live production incidents.
Monitor application performance, take steps to improve overall application performance and stability, and follow through with implementation.
Conduct system analysis, configuration management, and develop improvements for system software performance, availability, and reliability.
Design, write, ship, and motivate the creation of software and systems to increase observability, product reliability, and organizational efficiency.
Maintain and monitor deployment, orchestration, of servers,
docker containers, databases, and general backend infrastructure.
Develop Run Books/Standard Operating Procedure for recurring Production issues, also working on a permanent solve.
Perform Incident Analysis on a regular basis with the intention of preventing and finding a long-term solve for Incidents.
Educational Qualifications:
B Tech in Computer Science or related discipline preferred.
Key Skills:
Experience in monitoring and analyzing infrastructure performance using standard performance monitoring tools.
Demonstrable experience in Containerization-Docker and orchestration (Kubernetes).
Experience with Infrastructure As Code (Terraform, Cloud Formation, Ansible).
Knowledge and proven hands-on experience in large-scale databases and distributed technologies, such as Kafka and Confluent Platform Kafka.
Basic programming and scripting skills.
Experience Required:
Total Yrs of experience: 8-10
Major Stakeholders:
Internal
Product Manager from Digital Factory
Business Analyst from BTG team
Incident Management team
Development Team
Required Skills
1. Robust Linux and Networking Fundamentals
2. Programming concepts - Python is preferred
3. Jenkins, Terraform, Argo CD, AWS
4. Docker, Kubernetes, Podman, Helm
5. Familiarity with EKS Upgrades and pitfalls
6. CI/CD, SAST, SCA, Patch Management, Monitoring and ing
📌 Senior Site Reliability Engineer (Bengaluru)
🏢 HDFC Bank
📍 Bengaluru