07 Aug
|
programming.com
|
Bengaluru
07 Aug
programming.com
Bengaluru
Site Reliability Engineer (SRE) – Data Reliability & Platform Engineering
Location: Bangalore (Whitefield)
Work Mode: Work from Office (5 Days)
Experience: 5+ Years
Job Summary
We are looking for a skilled Site Reliability Engineer (SRE) to join our Platform Engineering team. In this role, you will ensure the reliability, scalability, and performance of cloud platforms and data pipelines. You will work primarily with Google Cloud Platform (GCP), focusing on infrastructure automation, observability, CI/CD, and cloud-native operations while improving system uptime and data reliability.
Key Responsibilities
Build and maintain highly available, scalable cloud infrastructure on GCP.
Develop and automate infrastructure using Terraform and Infrastructure as Code (IaC).
Design and manage CI/CD pipelines and GitOps workflows.
Implement monitoring, logging, and observability using Prometheus, Grafana, Cloud Monitoring, and OpenTelemetry.
Participate in incident response, root cause analysis (RCA), and production support.
Ensure reliability and performance of batch and streaming data pipelines.
Monitor data quality, freshness, lineage, and availability using modern data reliability practices.
Collaborate with platform, application,
and data engineering teams to improve cloud operations and automation.
Optimize cloud infrastructure, deployment processes, and operational efficiency.
Required Skills
5–8+ years of experience in Site Reliability Engineering, Platform Engineering, or Data Engineering.
Robust hands-on experience with Google Cloud Platform (GCP).
Expertise in Terraform, CI/CD, GitOps, and Infrastructure as Code.
Experience with Python, Go, Bash, or Shell scripting.
Strong knowledge of Docker, Kubernetes, and distributed systems.
Hands-on experience with Prometheus, Grafana, Cloud Monitoring, and OpenTelemetry.
Good understanding of Linux, networking, and cloud infrastructure.
Experience with BigQuery, Kafka, Spark, SQL/NoSQL databases, or modern data platforms.
Knowledge of Ansible and cloud automation tools is a plus.
Solid analytical, troubleshooting, and communication skills.
Preferred Skills
Exposure to OCI or Azure.
Experience with OpenLineage, data observability, and contemporary data reliability practices.
Experience working in Agile/DevOps environments.
Application Question(s):
Current CTC
Expected CTC
Currently serving notice period(Yes/ No) and mention the last working day?
Work Location: In person
📌 Site Reliability Engineer Sre – Data Reliability & Platform Engineering Bengaluru
🏢 programming.com
📍 Bengaluru