15 Sep
|
3across
|
Bengaluru
Senior Site Reliability Engineer (SRE)
Experience: 712 Years
Job Type: Full Time
We are looking for an experienced Site Reliability Engineer (SRE) to manage and improve the reliability, availability, and performance of critical production systems. The ideal candidate should have robust hands-on experience in SRE, Linux/Unix, Python and Shell scripting, along with experience in handling technical teams and critical production incidents.
Key Responsibilities
- Drive resolution and restoration of critical production incidents and ensure timely service recovery.
- Identify incident severity, business impact, risks, and appropriate escalation procedures.
- Troubleshoot complex issues across Linux/Unix-based production environments.
- Develop and maintain Python and Shell/Bash scripts for automation and operational efficiency.
- Perform root cause analysis (RCA) and contribute to Problem Management activities.
- Investigate and resolve application, infrastructure, and data-related production issues.
- Coordinate with cross-functional technology teams during critical incidents.
- Ensure proper incident documentation, timelines, impact analysis, and resolution details.
- Identify opportunities for automation, process improvement,
and operational efficiency.
- Support resiliency initiatives and contribute to improving system reliability.
- Conduct incident reviews and identify recurring/known issues.
- Handle and guide a team of technical professionals during production support activities.
- Ensure adherence to operational, audit, compliance, and change-management processes.
Required Skills
- Site Reliability Engineering (SRE)
- Python & Shell/Bash Scripting
- Linux / Unix
- Team Handling & Production Incident Management
Preferred Skills
- Major Incident Management
- Problem Management & RCA
- SQL
- Monitoring/Observability tools
- Automation
- ITIL / Service Management
- Cloud technologies
- Chaos Engineering / Resiliency
- ServiceNow or similar ITSM tools
Candidate Profile
- 7–12 years of relevant experience in SRE / Production Support / Application Support / Reliability Engineering.
- Strong hands-on troubleshooting and scripting capabilities.
- Proven experience handling critical production incidents.
- Experience in leading or mentoring technical teams.
- Strong communication, analytical, and problem-solving skills.
📌 Site Reliability Engineer (Bengaluru)
🏢 3across
📍 Bengaluru