Description
Job Responsibilities
- Own and improve the reliability, scalability, performance, and operational efficiency of critical OCI Compute services.
- Lead investigation and resolution of complex production incidents; drive mitigation, recovery, RCA, and follow-up improvements.
- Improve service KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures.
- Build automation and tooling to reduce recurring operational toil and improve production safety.
- Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements.
- Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response.
- Support major upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability management and production rollouts.
- Troubleshoot complex distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes.
- Contribute to incident-management practices, operational readiness, and service ownership improvements.
- Mentor team members and act as a technical resource for partner teams.
- Participate in a 12x7 on-call rotation and lead response during customer-impacting incidents.
Mandatory Skills
- Total Experience of 7+ Years
- 5+ years of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role.
- Experience operating and improving highly available distributed systems in production.
- Robust programming or scripting skills in Python, Java, Go, or similar languages.
- Strong hands-on experience with Linux, cloud infrastructure, networking, compute,
and storage.
- Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing.
- Experience defining or improving service SLIs, SLOs, KPIs, and operational procedures.
- Strong incident-management, troubleshooting, RCA, and problem-solving skills.
- Experience with deployment pipelines, release validation, automation, and change-management practices.
- Understanding of distributed systems, service dependencies, capacity planning, and performance tuning.
- Ability to work independently on complex technical issues and collaborate across engineering teams.
- Strong written and verbal communication skills.
Preferred Skills
- Experience with OCI and cloud infrastructure services.
- Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation.
- Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD.
- Experience with service migrations, fleet maintenance, upgrades, patching, or large-scale rollouts.
- Experience with architecture reviews, operational-readiness reviews, and post-incident improvements.
- Experience mentoring engineers or leading technical initiatives across teams.
- Familiarity with security, compliance, and access-control practices in production environments.
Self-Test Questions
1.
Does the candidate have 5+ years of SRE, Production Engineering, Cloud Operations, or Systems Engineering experience?
2. Has the candidate owned or significantly improved a critical production service or infrastructure component?
3. Can they lead a complex incident end-to-end: investigation, mitigation, recovery, RCA, and follow-up actions?
4. Do they have strong hands-on experience with Linux, cloud infrastructure, distributed systems, networking, compute, or storage?
5. Are they proficient in Python, Java, Go, or a similar language for automation, tooling, and troubleshooting?
6. Have they built or improved automation, CI/CD pipelines, deployment validation, operational tooling, or remediation workflows?
7. Do they have experience with monitoring, alerting, dashboards, logs, metrics, tracing, and SLO/KPI-driven reliability improvement?
8. Can they work independently on ambiguous technical problems, collaborate across engineering teams, mentor peers, and participate in a 12x7 on-call rotation?
Role Details
Field Requirement Role Site Reliability Engineer 4 (IC4) Experience 5+ Years Location Bangalore Only Work Mode Hybrid - 3 days from Office Primary Skills Site Reliability Engineering, Production Engineering, Linux, OCI/Cloud Infrastructure, Distributed Systems, Python/Java/Go, Automation, CI/CD, Monitoring and Observability, Incident Management, RCA, SLOs/SLIs/KPIs, Performance Tuning, Capacity Planning, AIOps, Deployment Validation, On-Call Operations, Security Vulnerability Management, Security, Vulnerability
35% Ops, 65% - Automation and Coding
Qualifications
Career Level - IC4
📌 Principal Site Reliability Engineer (Python, Java and Automation Specialization) (India)
🏢 Oracle
📍 India