- Handle SRE operational duties including responding to pull requests and maintaining smooth continuous integration and delivery (CI/CD) processes.
- Maintain and fine-tune applications for optimal performance and adherence to requirements.
- Explore and experiment with current technologies through Proof-of-Concepts to enhance functionalities or identify new opportunities.
- Automate deployment, configuration, and operational processes to improve efficiency and accuracy.
- Collaborate with development teams to guide system architecture, focusing on reliability, efficiency, and scalability.
- Implement and manage observability tools such as Grafana, Prometheus, and New Relic to monitor critical services effectively.
- Develop custom reliability tools and frameworks for engineering teams.
- Participate in on-call rotations for critical systems, lead incident responses, and conduct post-mortem analyses.
- Drive system and process efficiencies including capacity planning, configuration management, performance tuning, monitoring, and root cause analysis.
- Act as a consultant within the organization for best practices in infrastructure management and assist teams in effective infrastructure utilization.