- Operational Stability & SRE Practices
- Maintain high system availability through proactive monitoring and incident prevention
- Define and track SLIs/SLOs to measure service health and reliability
- Lead root cause analysis (RCA) and implement preventive fixes
- Automation & Tooling
- Build and maintain automation for repetitive operational tasks
- Improve deployment pipelines and operational workflows
- Develop scripts/tools to enhance productivity and reduce human intervention
- Service Design & Transition