**Job Description** **Job Responsibilities** + Improve the reliability, scalability, performance, and operational efficiency of assigned OCI Compute services and components. + Investigate and resolve complex production incidents; contribute to mitigation, recovery, RCA, and follow-up actions. + Own and improve service-level KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures for assigned systems. + Build automation and tooling to reduce operational toil and improve production safety. + Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements. + Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response. + Support upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability management and production rollouts. + Troubleshoot distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes. + Contribute to incident-management practices, operational readiness, and service ownership improvements. + Share technical knowledge and support team members through documentation, reviews, and collaboration. + Participate in a **12x7 on-call rotation** and support response to customer-impacting incidents. **Responsibilities** **Mandatory Skills** + **4-8 years** of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role. + Experience operating and improving highly available production systems. + Strong programming or scripting skills in **Python, Java, Go** , or similar languages. + Hands-on experience with Linux, cloud infrastructure, networking, compute, and storage. + Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing. + Experience owning or improving service **SLIs, SLOs, KPIs** , and operational procedures. + Solid incident troubleshooting, RCA, debugging, and problem-solving skills.
+ Experience with deployment pipelines, release validation, automation, and change-management practices. + Understanding of distributed systems, service dependencies, capacity planning, and performance tuning. + Ability to work independently on technical problems and collaborate effectively with engineering teams. + Strong written and verbal communication skills. **Preferred Skills** + Experience with **OCI** and cloud infrastructure services. + Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation. + Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD. + Experience with service migrations, fleet maintenance, upgrades, patching, or production rollouts. + Experience with architecture reviews, operational-readiness reviews, and post-incident improvements. + Experience contributing to technical initiatives, knowledge sharing, code reviews, or operational improvements within the team. + Familiarity with security, compliance, and access-control practices in production environments. **Self-Test Questions for TA** 1. Does the candidate have **4-8 years** of relevant SRE, Production Engineering, Cloud Operations, or Systems Engineering experience? 2. Has the candidate independently operated or improved a production service, system, or infrastructure component? 3. Can they investigate production incidents and contribute to mitigation, recovery, RCA, and follow-up actions? 4. Do they have hands-on experience with Linux, cloud infrastructure, distributed systems, networking, compute, or storage? 5. Are they proficient in Python, Java, Go,
or a similar language for automation, tooling, and troubleshooting? 6. Have they built or improved automation, deployment validation, CI/CD pipelines, or operational tooling? 7. Do they have experience with monitoring, alerting, logs, metrics, tracing, and service health indicators such as SLOs or KPIs? 8. Can they work independently on assigned technical problems, collaborate with partner teams, and participate in a **12x7 on-call rotation** ? Career Level - IC3 **About Us** Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs. We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing
[email protected] or by calling 1-(phone hidden) in the United States. Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
📌 Senior Site Reliability Engineer (Bengaluru)
🏢 Oracle
📍 Bengaluru