19 Aug
|
Infosys
|
Bengaluru
:
- The Site Reliability Engineering Lead is a senior hands on technical leader within the Corporate Technology and Operations organization
- This teammate is accountable for elevating the reliability resiliency and operational excellence of critical enterprise platforms across hybrid cloud and onprem environments
Key Responsibilities:
- Reliability Engineering Automation
- Architect and deliver automation solutions that eliminate toil reduce MTTR and increase service resilience
- Experience in Ansible Puppet or Chef is a plus
- Implement intelligent alerting anomaly detection and event correlation leveraging AI and AIOps tools
- Guide and enforce SLO SLI adoption across product teams ensuring metrics inform decision making and prioritization
- Utilize Infrastructure as Code IaC tools for automating deployment of assets within cloud tenants
- Observability Operational Excellence
- Ensure operational readiness of applications and platforms through resiliency testing chaos engineering and failure mode validation
- Cross Functional Leadership Influence
- Partner with Delivery Architecture Security and Risk teams to embed reliability and resilience into design and execution
- Standardization Documentation
- Develop maintain and enforce runbooks response playbooks and automated recovery patterns
- Follow best practices and internal processes for Non Functional requirements to improve resiliency and reliability
- Mentorship Technical Development
- Coach and mentor Associate Skilled and Senior SREs to build technical depth and operational discipline
- Provide thought leadership in SRE methodologies cloud native operational patterns and automated reliability engineering
- Incident Leadership Production Operations
- Lead P1 P0 incident bridges and direct technical investigation efforts
- Perform hands on triage using logs traces metrics and application telemetry
- Drive mitigation recovery RCA development and follow through remediation
- Provide executive communications during major incidents
- Build operational automation based on recurring production issues
- Establish credibility through technical leadership during live service disruptions
Technical Requirements:
- 7 years of experience in Site Reliability Engineering DevOps Platform Engineering or Infrastructure Operations
- Deep hands on experience with distributed systems container orchestration Kubernetes and cloud native operational tooling
- Proficiency with automation and scripting languages Python Go PowerShell Ansible
- Strong understanding of observability platforms Splunk Dynatrace and event driven monitoring
- Proven leadership in major incident management and cross team technical coordination
- Strong grasp of networking Linux Unix internals and modern infrastructure patterns
- Excellent communication skills including executive level situational awareness during critical incidents
- Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices
Additional Responsibilities:
- Experience enabling large scale SRE transformations or modernization initiatives
- Demonstrated proficiency with GitLab Duo or similar AI technologies
- Familiarity with chaos engineering resilience assessments and service failure modeling
- Exposure to hybrid cloud and multi cloud operational frameworks
- Experience contributing to or leading Center for Enablement functions or Communities of Practice
- Expertise with highly regulated industries preferred
Preferred Skills:
Technology->Container Platform->Kubernetes,Technology->Devops->Ansible,Technology->DevOps->Site Reliability Engineering(SRE),Technology->OpenSystem->Python - OpenSystem
📌 Site Reliability Engineering Lead_Truist (Bengaluru)
🏢 Infosys
📍 Bengaluru