18 Aug
|
Infosys
|
Bengaluru
7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations. Deep hands‑on experience with distributed systems, container orchestration (Kubernetes), and cloud-native operational tooling. Proficiency with automation and scripting languages (Python, Go, PowerShell, Ansible).
Strong understanding of observability platforms (Splunk, Dynatrace) and event-driven monitoring. Proven leadership in major incident management and cross-team technical coordination. Robust grasp of networking, Linux/Unix internals, and modern infrastructure patterns.
Excellent communication skills, including executive-level situational awareness during critical incidents. Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices.
Reliability Engineering &
- Automation Architect and deliver automation solutions that eliminate toil, reduce MTTR, and increase service resilience.
Experience in Ansible, Puppet or Chef is a plus. Implement intelligent alerting, anomaly detection, and event correlation leveraging AI and AIOps tools. Guide and enforce SLO/SLI adoption across product teams, ensuring metrics inform decision-making and prioritization.
Utilize
Infrastructure-as-Code (IaC) tools for automating deployment of assets within cloud tenants. Observability &
- Operational Excellence Ensure operational readiness of applications and platforms through resiliency testing, chaos engineering, and failure-mode validation. Cross-Functional Leadership &
- Influence Partner with Delivery, Architecture, Security,
and Risk teams to embed reliability and resilience into design and execution. Standardization &
- Documentation Develop, maintain, and enforce runbooks, response playbooks, and automated recovery patterns. Follow best practices and internal processes for Non-Functional requirements to improve resiliency and reliability. Mentorship &
- Technical Development Coach and mentor Associate, Professional, and Senior SREs to build technical depth and operational discipline. Provide thought leadership in SRE methodologies, cloud-native operational patterns, and automated reliability engineering.
Incident
Leadership &
- Production Operations Lead P1/P0 incident bridges and direct technical investigation efforts. Perform hands-on triage using logs, traces, metrics, and application telemetry. Drive mitigation, recovery, RCA development, and follow-through remediation. Provide executive communications during major incidents. Build operational automation based on recurring production issues. Establish credibility through technical leadership during live service disruptions.
Experience enabling large-scale SRE transformations or modernization initiatives. Demonstrated proficiency with GitLab Duo, or similar AI technologies. Familiarity with chaos engineering, resilience assessments, and service failure modeling. Exposure to hybrid-cloud and multi-cloud operational frameworks.
Experience contributing to or leading Center for Enablement functions or Communities of Practice. Expertise with highly regulated industries preferred.
📌 Site Reliability Engineering Lead_Truist (Bengaluru)
🏢 Infosys
📍 Bengaluru