We are looking for a Site Reliability Engineer to join our team and ensure the continuous, reliable operation of company services. This role involves proactive monitoring, incident response, and building resilient observability and escalation practices across our infrastructure.
What You ll Do
- Ensure monitoring and uninterrupted operation of company services.
- Write and maintain alerting rules and runbooks.
- Perform triage of incoming incidents and initial diagnosis of issues.
- Build and maintain escalation chains for incident response.
- Perform technical incidents resolution activities according to runbooks.
- Participate in on-call rotations and post-incident reviews (RCA/postmortems).
- Continuously improve observability coverage and reduce alert noise/false positives.
- Collaborate with development and infrastructure teams to identify reliability risks and implement preventive measures.
What You ll Bring
- Experience with observability tools (Grafana, ELK, VictoriaMetrics).
- Experience working with Linux.
- Experience working with Kubernetes (k8s).
- Experience with AWS and Azure cloud platforms.
- Ability to analyze incidents, identify root causes, and propose remediation steps.
Bonus Skills
- Experience with Infrastructure as Code (Terraform, Ansible, or similar).
- Scripting skills (Python, Bash) for automation of operational tasks.
- Understanding of DevOps and CI/CD principles.
- Experience with incident management tools (PagerDuty, Opsgenie, etc.).
- Effective communication skills and a cooperative approach to teamwork.
📌 System Reliability Engineer III (Bengaluru)
🏢 Securiti.Ai
📍 Bengaluru
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.