Job Requirements
- Drive
operational stability
across ERP and other finance applications
- Implement and enhance
automation and tooling
to reduce manual effort and improve efficiency
- Own and execute
Disaster Recovery (DR) planning and testing
to ensure business continuity
- Lead
service design and service transition
activities for new and existing systems
- Manage
incident, problem, and change processes
aligned with SRE and ITIL practices
- Establish effective
service communication frameworks
for incidents and outages
- Drive
continual service improvement (CSI)
initiatives across finance systems
- Define and manage
service metrics, SLAs, SLIs, and reporting dashboards
- Collaborate with engineering teams to improve
system reliability, observability, and performance
- Ensure smooth onboarding and ownership of
satellite applications within the finance ecosystem
- Proactively identify risks and implement preventive measures to minimize production issues
Preferred Skills and Experience
- Solid experience in Site Reliability Engineering / Production Support / DevOps roles
- Experience supporting Oracle Cloud ERP or similar enterprise SaaS platforms
- Knowledge of monitoring, alerting, and observability tools (e.g., Splunk, Grafana, OCI monitoring)
- Experience in incident management, RCA, and problem management
- Exposure to automation frameworks and scripting (e.g., Python, Shell, Terraform)
- Understanding of cloud platforms (OCI/AWS/Azure) and distributed systems
- Familiarity with ITIL processes and service management frameworks
- Experience working in global, distributed teams
- Exposure to financial systems and processes is an advantage
Key Responsibilities:
- Operational Stability & SRE Practices
- Maintain high system availability through proactive monitoring and incident prevention
- Define and track SLIs/SLOs to measure service health and reliability
- Lead root cause analysis (RCA) and implement pr