Job Requirements The Site Reliability Engineer (P2) ensures the reliability, availability, and performance of software systems through monitoring, automation, and incident response.
Key Responsibilities
- System Reliability & Performance
- Monitor and maintain system availability, performance, and capacity using observability tools
- Implement monitoring solutions and alerts to proactively identify potential issues
- Automation & Infrastructure
- Develop and maintain automation scripts to reduce manual operational tasks
- Implement infrastructure as code solutions using established tools and frameworks
- CI/CD Pipeline Support
- Support continuous integration and deployment pipelines to enable accelerated software delivery
- Collaborate with development teams to optimize build and deployment processes
- Incident Management
- Respond to and resolve incidents following established procedures and best practices
- Participate in on-call rotations and conduct post-incident analysis to prevent recurrence
- Problem Solving
- Analyze and solve variety of problems using experience, judgment, and past precedents
- Apply knowledge of system architecture to troubleshoot and resolve technical issues
- Collaboration
- Work with development and operations teams to improve system reliability and performance
- Receive moderate guidance on complex technical challenges while providing informal guidance to newer team members
Requirements 2+ years of programming experience. This is not a programming role, but often there will be tasks that require bash or Python scripts. You should be comfortable creating easy scripts and performing general UNIX system administration tasks. Experience in technology integration and analysis, using common technologies like SAML, LDAP, etc.
Experience with modern JavaScript technologies, ReactJS/EmberJS, NodeJS.