We are looking for a Cloud/SRE professional to maintain platform reliability, performance, monitoring, scalability and incident response.
Key Responsibilities
- Define and monitor SLOs, error budgets and uptime targets.
- Build monitoring dashboards and alerts.
- Handle production incidents and root-cause analysis.
- Plan autoscaling and capacity requirements.
- Perform load testing and reliability improvements.
- Optimise cloud infrastructure costs.
- Improve system resilience and fault tolerance.
- Maintain incident runbooks.
- Work with backend teams on performance issues.
- Conduct disaster recovery drills.
Required Skills
- Experience in SRE, cloud operations or reliability engineering.
- Solid AWS/GCP/Azure knowledge.
- Monitoring and observability experience.
- Real-world incident management/on-call experience.
- Python, Go or Bash scripting.
- Understanding of distributed systems and caching.
- Linux troubleshooting and performance tuning.
- Which cloud platform are you strongest in?
- Do you have production on-call experience?
- Which monitoring tools have you used?
- Have you worked with Kubernetes?
- Have you handled production incidents and root-cause analysis?
Experience:
- Site Reliability Engineer: 4 years (Required)
Location:
- Greater Noida, Uttar Pradesh (Required)
Work Location: In person
📌 Cloud / Site Reliability Engineer (SRE) (India)
🏢 NCR BACKUP ACADEMY
📍 India
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.