12 Sep
|
3across
|
Bengaluru
SRE Lead
Exp: 7-12 Yrs
Location: Bangalore
We are looking for an experienced Site Reliability Engineer (SRE) with strong expertise in Python scripting, Linux/Unix, SQL, production support, incident management, and automation. The ideal candidate should have hands-on experience working in L3/L4 support environments, managing critical production incidents, driving root-cause analysis, and leading/handling support teams.
Key Responsibilities
- Drive and coordinate L3/L4 production support activities for critical business applications and services.
- Own the end-to-end incident management and service restoration process for high-impact incidents.
- Assess incident severity, priority, business impact, customer impact, and risk and ensure appropriate escalation.
- Lead troubleshooting and restoration activities for Sev1Sev4 incidents.
- Coordinate with application, infrastructure, database, network, cloud, and other technology teams to restore services within agreed SLAs.
- Perform detailed root-cause analysis (RCA) and drive corrective and preventive actions.
- Utilize Python scripting for automation, troubleshooting, monitoring, operational improvements, and repetitive task reduction.
- Develop and maintain scripts/tools to automate incident resolution, health checks, operational activities, and service recovery processes.
- Troubleshoot production issues using Linux/Unix commands, SQL queries, logs, monitoring tools, and system/application data.
- Investigate and remediate customer/client data issues and coordinate with relevant technology teams for resolution.
- Perform activities such as batch restarts, service restarts, routing changes, contingency procedures, and controlled recovery actions as required.
- Engage with external software/hardware vendors when specialized technical support is required.
- Drive High Impact Incident Communications, providing timely updates on incident status,
business/customer impact, troubleshooting progress, and service restoration.
- Ensure incident tickets contain accurate and complete information, including impact, timeline, actions taken, participants, resolution, and RCA details.
- Work closely with Problem Management teams to provide detailed incident information and support creation of Problem Records.
- Identify and document known errors, repeatable incidents, and recurring production issues in the Known Error Database.
- Conduct regular incident reviews to identify trends, recurring issues, and opportunities for service improvement.
- Create and maintain standard operating procedures, technical documentation, troubleshooting guides, and incident playbooks.
- Identify opportunities for task automation, tooling improvements, and operational efficiency.
- Participate in resiliency exercises and Chaos Engineering activities to improve application and infrastructure reliability.
- Support audit, compliance, and risk remediation activities related to production operations.
- Participate in high-risk and complex technology changes, including Permit-to-Operate / change governance activities.
- Monitor service reliability and proactively identify potential production risks and failure points.
- Ensure appropriate access and temporary access requests are reviewed and managed in accordance with defined procedures.
- Lead or coordinate support teams during critical incidents and ensure effective collaboration across technology functions.
- Mentor and guide team members on incident management, troubleshooting, production support, automation, and SRE best practices.
- Ensure the support team follows defined processes, SLAs, escalation procedures, and operational standards.
- Drive continuous improvement initiatives to increase service availability, reliability, automation, and operational efficiency.
Mandatory Skills
- 7+ years of overall experience in Site Reliability Engineering / Production Support / Application Support.
- Strong hands-on experience in Python scripting.
- Proven experience in L3/L4 production support.
- Strong knowledge of Linux/Unix environments.
- Robust hands-on experience with SQL and database troubleshooting.
- Experience in incident management and major/high-impact incident handling.
- Strong experience in RCA, problem management, and service restoration.
- Experience with production monitoring, troubleshooting, log analysis, and operational tools.
- Experience in automation and scripting to reduce manual operational activities.
- Proven experience in handling/leading support teams and coordinating multiple technology teams during critical incidents.
- Strong understanding of SLA, incident severity, escalation, change management, and IT service management processes.
- Excellent communication and stakeholder management skills.
- Ability to work under pressure and take ownership during critical production incidents.
Good to Have
- Experience with SRE/DevOps practices and tools.
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Experience with monitoring and observability tools.
- Knowledge of CI/CD and deployment processes.
- Exposure to Chaos Engineering and resiliency testing.
- ITIL / relevant industry certification.
- Experience supporting large-scale, business-critical applications.
📌 SRE Lead (Bengaluru)
🏢 3across
📍 Bengaluru