The Site Reliability Engineer will improve reliability, observability, service availability, incident response and root cause analysis for AWS-hosted systems.
- Amazon CloudWatch, AWS X-Ray, logs, metrics and dashboarding
- Incident management, RCA and problem management practices
- Scripting and automation using Python, shell or related tools
- Availability, resilience, performance and capacity concepts
Key Responsibilities
- Monitor application and infrastructure availability, performance and reliability.
- Build dashboards, alerts and service health metrics.
- Support incidents, RCA and corrective action tracking.
- Automate operational and reliability improvement tasks.
- Partner with DevOps,
application and operations teams to improve resilience.
- Monitoring dashboards and alert rules
- Incident and RCA reports
- Reliability improvement backlog
- Operational automation scripts
Qualification
- B.E. / B.Tech / M.Tech / MCA or equivalent degree in Computer Science, Information Technology, Engineering or related discipline.
- Hands-on delivery experience aligned to the stated experience range and role scope.
Good to have
- AWS certification or active AWS certification plan.
- Experience in reusable assets, accelerators, cloud competency building, customer workshops or proposal support.
📌 Site Reliability Engineer (Hyderabad)
🏢 Sonata Software
📍 Hyderabad
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.