Solve complex problems related to infrastructure cloud services and build automation to prevent problem recurrence. Design, write, and deploy software to improve the availability, scalability, and efficiency of Oracle products and services. Design and develop designs, architectures, standards, and methods for large-scale distributed systems. Facilitate service capacity planning and demand forecasting, software performance analysis, and system tuning.
As a Site Reliability Engineer, you will work with the Production Engineering and SRE teams to own, run, and improve critical healthcare services. You will help keep our cloud-native EHR platforms reliable, secure, scalable, and easy to operate.
You will understand how services are built, deployed, monitored, and supported in production. You will work closely with development teams to improve service design, reduce failures, automate manual work, and improve performance.
You will also help use AI and AIOps to improve operations, including smarter alerting, faster incident detection, automated troubleshooting,
and better root cause analysis.
Key Responsibilities
Own the reliability, availability, performance, and operations of production services.
Support cloud-native EHR platforms built with microservices, Kubernetes, and OCI.
Understand service architecture, dependencies, capacity, security, and failure points.
Improve monitoring, alerting, observability, and incident response.
Use AI, automation, and AIOps to reduce manual work and improve system health.
Build tools and scripts for deployment, monitoring, recovery, and operational tasks.
Troubleshoot complex production issues and drive them to resolution.
Lead root cause analysis for major incidents and help prevent repeat issues.
Partner with development teams to improve service design and operability.
Create and maintain SOPs, runbooks, dashboards, and knowledge articles.
Support migration and modernization of existing hosting
📌 Site Reliability Developer Noida (India)
🏢 Oracle
📍 India