- Generating response
- Copilot said
- Site Reliability Engineer SRE Roles Responsibilities
- Monitor application and infrastructure health 24 7
- Ensure high availability reliability and performance of systems
- Respond to incidents outages and production issues
- Perform root cause analysis RCA and implement permanent fixes
- Automate repetitive operational and deployment tasks
- Develop and maintain monitoring alerting and logging solutions
- Configure and manage CI CD pipelines
- Support application deployments and release management
- Manage cloud infrastructure Azure AWS or GCP
- Perform capacity planning and resource optimization
- Implement Infrastructure as Code IaC using tools like Terraform or ARM templates
- Collaborate with development teams to improve application reliability
- Define and track SLAs SLIs and SLOs
- Manage backups disaster recovery and business continuity plans
- Ensure security compliance and vulnerability remediation
- Troubleshoot Linux networking database and application issues
- Maintain Kubernetes and containerized environments
- Create operational documentation and runbooks
- Conduct post incident reviews and recommend improvements
- Continuously improve system scalability efficiency and stability
- Key Skills
- Linux Unix Administration
- Azure AWS GCP
- Kubernetes Docker
- Python Bash PowerShell
- Terraform
- Jenkins Azure DevOps GitHub Actions
- Prometheus Grafana Splunk Azure Monitor
- Networking Security
- Incident Management and Problem Solving