03 Sep
|
Xebia IT Architects
|
Bengaluru
03 Sep
Xebia IT Architects
Bengaluru
SRE - Role purpose The L2/SRE team investigates and resolves incidents end-to-end as a resolver, not a router, and applies reliability-engineering practices.
Key responsibilities
- Investigate and resolve P2-P4 incidents through to closure, taking full operational ownership.
- Perform root-cause analysis and drive structured problem management to prevent recurrence.
- Convert recurring incidents and toil into automations and runbooks, working with the automation factory.
- Apply SLO and error-budget practices; monitor service health and act proactively.
- Support major-incident response under MIM coordination; participate in on-call rotation.
- Contribute to detection tuning, knowledge articles and continuous service improvement.
Experience & skills 1-6 years in SRE or operations engineering; strong Linux and cloud fundamentals; scripting ability in at least one of Python or Bash (or equivalent); familiarity with observability and incident tooling; solid troubleshooting and problem-management discipline.
Core technical stack:
- SRE / DevOps background
- AWS solid hands-on experience
- Docker & Kubernetes
- Terraform / IaC
- Python or Bash scripting
- CI/CD pipelines
- Monitoring / Observability / SRE tools
- Strong production troubleshooting
📌 SRE Engineer (Bengaluru)
🏢 Xebia IT Architects
📍 Bengaluru