Gritstone Technologies is a rapid growing software and technology company dedicated to building reliable, scalable, and innovative solutions for businesses. We are looking for a skilled and proactive Site Reliability Engineer to help us maintain and improve the performance, availability, and reliability of our systems and infrastructure.
Monitor, maintain, and improve the reliability, availability, and performance of production systems and services. Define and track Service Level Objectives, Service Level Indicators, and error budgets. Identify and resolve infrastructure and application bottlenecks, outages, and performance issues.
Design and implement automation to reduce manual operational work and improve system efficiency. Build and maintain CI/CD pipelines,
deployment processes, and infrastructure as code. Conduct root cause analysis for incidents and implement preventive measures to avoid recurrence.
Collaborate with development teams to ensure reliability and scalability are built into new features and services from the ground up. Manage and optimize cloud infrastructure across platforms such as AWS, GCP, or Azure. Maintain documentation for systems, runbooks, and operational procedures.
Participate in on-call rotations and incident response processes.