24 Sep
|
Orcapod Consulting Services
|
Bengaluru
24 Sep
Orcapod Consulting Services
Bengaluru
We are looking for an experienced Service Reliability Engineer (SRE) with strong expertise in Production Support, Application Support, DevOps, Enterprise SaaS/ERP, Cloud, Monitoring, Incident Management, and Automation.
The candidate will be responsible for ensuring the reliability, availability, performance, and scalability of critical enterprise applications in a 24x7/global production environment. Experience with Oracle Cloud ERP / Oracle Fusion Cloud is preferred.
Key Responsibilities
1. Production Support & SRE
- Provide L2/L3 production support for enterprise applications and SaaS/ERP platforms.
- Monitor application health, availability, performance, and capacity.
- Identify and resolve production issues within defined SLAs.
- Participate in on-call and 24x7 support rotations.
- Drive reliability improvements and operational excellence.
1. Enterprise SaaS / ERP Support
- Support enterprise SaaS and ERP applications in production environments.
- Experience with Oracle Cloud ERP / Oracle Fusion Cloud is preferred.
- Troubleshoot application, integration, performance, and availability issues.
- Coordinate with application, infrastructure, cloud, and development teams.
1. Monitoring & Observability
- Monitor applications and infrastructure using tools such as:
- Splunk
- Grafana
- OCI Monitoring
- Prometheus
- AWS CloudWatch
- Develop and maintain dashboards, alerts, and monitoring solutions.
- Analyze logs, metrics, and application health indicators.
1. Incident & Problem Management
- Handle P1/P2 production incidents and participate in incident response.
- Perform detailed Root Cause Analysis (RCA) for recurring and critical incidents.
- Implement preventive and corrective actions.
- Participate in post-incident reviews and continuous improvement initiatives.
- Ensure incidents are resolved within agreed service levels.
1. Cloud & Infrastructure
- Work with cloud platforms such as OCI, AWS, or Azure.
- Troubleshoot cloud infrastructure and application-related issues.
- Monitor cloud resources, availability, and performance.
- Support cloud-based enterprise applications and services.
1. Automation & Scripting
- Develop automation solutions using:
- Python
- Shell scripting
- Terraform
- Automate repetitive operational and support activities.
- Contribute to infrastructure and application automation initiatives.
1. Service Management
- Follow IT service management processes covering:
- Incident Management
- Change Management
- Problem Management
- Apply ITIL principles in production support activities.
- Maintain proper operational documentation and support procedures.
1. Reliability Engineering
- Define and monitor SLI, SLO, and SLA metrics.
- Improve application availability, performance, scalability, and resilience.
- Identify reliability risks and recommend appropriate improvements.
- Contribute to capacity and performance management.
1. Disaster Recovery
- Participate in Disaster Recovery (DR) planning and testing.
- Understand and work with RTO and RPO requirements.
- Support application recovery and business continuity exercises.
1. Stakeholder Management
- Work closely with global teams, application owners, development teams, infrastructure teams, and business stakeholders.
- Provide clear communication during critical incidents and outages.
- Coordinate effectively in a 24x7/global support environment.
Must-Have Skills
- Experience in SRE / Production Support / Application Support / DevOps.
- Experience supporting Enterprise SaaS / ERP applications in production.
- Experience with Oracle Cloud ERP / Oracle Fusion Cloud is preferred.
- Strong experience with monitoring and observability tools such as Splunk, Grafana, OCI Monitoring, Prometheus, or CloudWatch.
- Solid P1/P2 incident management and RCA experience.
- Experience with OCI, AWS, or Azure.
- Strong scripting/automation skills using Python, Shell, or Terraform.
- Knowledge of Incident, Change, and Problem Management.
- Exposure to ITIL processes.
- Understanding of SLI, SLO, SLA, availability, performance, and scalability.
- Experience with DR planning/testing, RTO, and RPO.
- Strong stakeholder management and communication skills.
- Experience working in 24x7/global production support environments.
📌 Application Support Engineer (Bengaluru)
🏢 Orcapod Consulting Services
📍 Bengaluru