We are hiring a Site Reliability Engineer who can do three things well: take a production failure down to its actual root cause and close it : including when the cause turns out to be the alert itself rather than the service; build the automation that stops that class of failure from ever needing a human again; and own a project end-to-end : taking a service onto our release tooling, standing up a new tenant's queues, promoting a set of SLOs : from scope to production.
We are a fundamentally proactive team. We care about release process and tooling because that's how we prevent outages before they happen. We invest in tooling because tooling : not heroics : is how we stop the next incident. This isn't a reactive ops role; it's a role for someone who treats every incident as an unfinished automation task.
7-9 years of experience in SRE, DevOps, or infrastructure/production engineering roles
Strong track record of driving production issues to true root cause :
not just mitigating symptoms
Proven ability to build automation (scripts, tools, services) that durably eliminates a class of manual or repetitive work
Demonstrated experience owning technical projects independently from scoping through to production delivery
Strong Python skills : comfortable building and maintaining automation, tooling, and scripts used by other engineers
Advanced skills reading and correlating logs, metrics, and traces to diagnose issues across multiple interdependent systems
Extensive experience with release engineering, deployment tooling, or CI/CD pipelines
Excellent written communication : explicit incident writeups, project scoping docs, and technical updates that influence direction across teams
Experience mentoring engineers and leading technical initiative.
📌 SRE Lead (India)
🏢 Artech Infosystems Private
📍 India
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.