Key Responsibilities
- Monitor and maintain the operational health of critical services.
- Work closely with development teams to improve service reliability and performance.
- Manage production incidents, troubleshoot issues, and perform root cause analysis (RCA).
- Automate repetitive operational tasks using Python, Bash, or JavaScript.
- Build and enhance monitoring dashboards, KPIs, alerts, and alarming systems.
- Develop runbooks and automate operational processes to reduce manual effort.
- Scale systems through automation and infrastructure improvements.
- Drive incident response, reduce operational toil, and improve system resilience.
- Manage ticket prioritization and ensure timely issue resolution.
- Support CI/CD pipelines and deployment automation.
Mandatory Skills
Linux Administration (SysAdmin)
Linux/Unix Commands
Python / Bash / JavaScript Scripting
CI/CD Pipelines
Automation & Operational Tooling
Incident Management & Root Cause Analysis
Monitoring, Alerting & Dashboard Development
Service Reliability & Performance Optimization
Production Support
📌 Site Reliability Engineer (SRE) (India)
🏢 Cerebra
📍 India