23 Aug
|
Aziro
|
Hyderabad
Job Title: Site Reliability Engineer (SRE) - Google Cloud Platform
About Us
We are a forward-thinking organization looking to enhance our infrastructure and operational efficiency. As we grow, we are seeking a talented Site Reliability Engineer (SRE) with expertise in Google Cloud Platform (GCP) to optimize, automate, and manage the deployment of our cloud-based infrastructure.
Role Overview
As an SRE on our team, you will play a key role in ensuring the reliability, performance, and scalability of our systems. You will monitor, optimize, and automate the deployment of our infrastructure using GCP services, including Cloud Run, Pub/Sub, Virtual Machines, Kubernetes, and Dataflow. You will also take ownership of our logging and monitoring tools, streamline our deployment pipelines, and work to enhance database performance.
Key Responsibilities
- Infrastructure Management:
- Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform (GCP).
- Work with GCP services like Cloud Run, Pub/Sub, Virtual Machines, Kubernetes Engine, App Engine, and Dataflow.
- Automate and optimize deployment processes using Infrastructure as Code (IaC) preferably Terraform and CI/CD pipelines.
- Monitoring and Optimization:
- Enhance system observability by setting up and maintaining logging and monitoring solutions, including Sentry.
- Proactively monitor the health and performance of our systems, troubleshoot incidents, and implement resolutions.
- Continuously identify opportunities to improve system efficiency and reduce latency.
- Deployment Pipeline Management:
- Oversee the deployment pipeline using Bitbucket to ensure smooth, reliable, and repeatable deployment processes.
- Collaborate with development teams to ensure the integration of scalable and maintainable deployment practices.
- Database Optimization:
- Monitor and maintain MySQL databases, focusing on performance and reliability.
- Analyze and optimize database indexes and address poorly performing queries.
- Collaboration and Best Practices:
- Work closely with software developers, QA engineers, and product managers to deliver high-quality, reliable solutions.
- Advocate for best practices in SRE, including monitoring, alerting, and incident management.
Qualifications:
- Proven experience as a Site Reliability Engineer or similar role.
- Hands-on expertise with Google Cloud Platform (GCP), including Pub/Sub, Virtual Machines, Kubernetes Engine, App Engine, and Dataflow.
- Strong scripting and automation skills using tools like Terraform, Python, bash, or equivalent.
- Proficiency in setting up and managing alerting, monitoring and logging tools (e.g., Sentry, Stackdriver, or similar) as well as dashboarding (e.g. Grafana).
- Experience managing CI/CD pipelines with Bitbucket or similar tools.
- Solid understanding of MySQL database performance tuning and query optimization.
- Familiarity with containerization and orchestration tools, such as Docker and Kubernetes.
Preferred Skills:
- Experience with Infrastructure as Code (IaC) using Terraform or similar tools.
- Knowledge of incident response and on-call best practices.
- Familiarity with up-to-date software delivery practices like DevOps and Agile.
📌 Site Reliability Engineer (SRE) - Google Cloud Platform (Hyderabad)
🏢 Aziro
📍 Hyderabad