SRE Lead (Bengaluru)

SRE Lead (Bengaluru)

12 Sep
|
3across
|
Bengaluru

12 Sep

3across

Bengaluru

SRE Lead

Exp: 7-12 Yrs

Location: Bangalore

We are looking for an experienced Site Reliability Engineer (SRE) with strong expertise in Python scripting, Linux/Unix, SQL, production support, incident management, and automation. The ideal candidate should have hands-on experience working in L3/L4 support environments, managing critical production incidents, driving root-cause analysis, and leading/handling support teams.

Key Responsibilities

- Drive and coordinate L3/L4 production support activities for critical business applications and services.
- Own the end-to-end incident management and service restoration process for high-impact incidents.
- Assess incident severity, priority, business impact, customer impact, and risk and ensure appropriate escalation.
- Lead troubleshooting and restoration activities for Sev1Sev4 incidents.
- Coordinate with application, infrastructure, database, network, cloud, and other technology teams to restore services within agreed SLAs.
- Perform detailed root-cause analysis (RCA) and drive corrective and preventive actions.
- Utilize Python scripting for automation, troubleshooting, monitoring, operational improvements, and repetitive task reduction.
- Develop and maintain scripts/tools to automate incident resolution, health checks, operational activities, and service recovery processes.
- Troubleshoot production issues using Linux/Unix commands, SQL queries, logs, monitoring tools, and system/application data.
- Investigate and remediate customer/client data issues and coordinate with relevant technology teams for resolution.
- Perform activities such as batch restarts, service restarts, routing changes, contingency procedures, and controlled recovery actions as required.
- Engage with external software/hardware vendors when specialized technical support is required.
- Drive High Impact Incident Communications, providing timely updates on incident status,



business/customer impact, troubleshooting progress, and service restoration.
- Ensure incident tickets contain accurate and complete information, including impact, timeline, actions taken, participants, resolution, and RCA details.
- Work closely with Problem Management teams to provide detailed incident information and support creation of Problem Records.
- Identify and document known errors, repeatable incidents, and recurring production issues in the Known Error Database.
- Conduct regular incident reviews to identify trends, recurring issues, and opportunities for service improvement.
- Create and maintain standard operating procedures, technical documentation, troubleshooting guides, and incident playbooks.
- Identify opportunities for task automation, tooling improvements, and operational efficiency.
- Participate in resiliency exercises and Chaos Engineering activities to improve application and infrastructure reliability.
- Support audit, compliance, and risk remediation activities related to production operations.
- Participate in high-risk and complex technology changes, including Permit-to-Operate / change governance activities.
- Monitor service reliability and proactively identify potential production risks and failure points.
- Ensure appropriate access and temporary access requests are reviewed and managed in accordance with defined procedures.
- Lead or coordinate support teams during critical incidents and ensure effective collaboration across technology functions.




- Mentor and guide team members on incident management, troubleshooting, production support, automation, and SRE best practices.
- Ensure the support team follows defined processes, SLAs, escalation procedures, and operational standards.
- Drive continuous improvement initiatives to increase service availability, reliability, automation, and operational efficiency.

Mandatory Skills

- 7+ years of overall experience in Site Reliability Engineering / Production Support / Application Support.
- Strong hands-on experience in Python scripting.
- Proven experience in L3/L4 production support.
- Strong knowledge of Linux/Unix environments.
- Robust hands-on experience with SQL and database troubleshooting.
- Experience in incident management and major/high-impact incident handling.
- Strong experience in RCA, problem management, and service restoration.
- Experience with production monitoring, troubleshooting, log analysis, and operational tools.
- Experience in automation and scripting to reduce manual operational activities.
- Proven experience in handling/leading support teams and coordinating multiple technology teams during critical incidents.
- Strong understanding of SLA, incident severity, escalation, change management, and IT service management processes.
- Excellent communication and stakeholder management skills.
- Ability to work under pressure and take ownership during critical production incidents.

Good to Have

- Experience with SRE/DevOps practices and tools.
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Experience with monitoring and observability tools.
- Knowledge of CI/CD and deployment processes.
- Exposure to Chaos Engineering and resiliency testing.
- ITIL / relevant industry certification.
- Experience supporting large-scale, business-critical applications.

📌 SRE Lead (Bengaluru)
🏢 3across
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: sre lead (bengaluru) / bengaluru