Site Reliability Engineer (SRE) – GCP Platform (Bengaluru)

Site Reliability Engineer (SRE) – GCP Platform (Bengaluru)

19 Aug
|
ITC Infotech
|
Bengaluru

19 Aug

ITC Infotech

Bengaluru

Role - Site Reliability Engineer (SRE) – GCP Platform

Location - Bangalore (O Shaughnessy Road)

Work type - Work from Office

Experience - 8+ Years

Key Responsibilities

Reliability & Operations

- Own end-to-end production systems reliability, availability, scalability, cost and performance.
- Drive measurable improvements in MTTR, MTTA, and incident response practices using automation and runbook additions and process enhancements.
- Participate in 24x7 on-call rotations and handle high-severity incidents and document the learnings on ongoing basis.
- Establish and manage SLI, SLO, SLA, Error Budgets, and operational metrics for mission critical services and partner with engineering teams with full accountability for upholding the SLOs.
- Partner with the various engineering, operations and cloud management teams to deliver highly reliable service in a timely manner.

Cloud & Infrastructure

- Design, deploy, and manage infrastructure on Google Cloud Platform (GCP).
- Work extensively on:
- GKE (Kubernetes Engine)
- Compute, networking, IAM, Load Balancers, TLS Certs
- BigQuery, Pub/Sub, cloud logging enhancement, metrics and logs analysis
- Implement and manage infrastructure using Terraform (Infrastructure as Code).

Kubernetes & Containers

- Deploy and manage containerized workloads using Kubernetes (GKE).
- Troubleshoot issues related to:
- Pods, nodes, networking, storage, services on an ongoing basis
- Manage deployments using Helm, YAML, and rollout strategies (Canary/Blue-Green).

Automation & CI/CD

- Build and maintain CI/CD pipelines using:
- Jenkins (pipeline-based, Groovy / Shell / Python scripting)
- Strong experience in using GitHub as a PowerUser
- Develop automation using Python and Shell scripting.
- Reduce operational toil through automation initiatives.





Observability & Monitoring

- Implement and manage monitoring systems using:
- Dynatrace, Grafana, logs and metrics explorer
- Work with logs, metrics, and traces for deep observability to identify trends and arrest problems proactively.
- Define alerting strategies based on system behaviour and SLOs and create runbooks.
- Work alongside operations teams to identify, fix the production incidents and own the problem resolution.
- Work with engineering teams to isolate infra and application issues and set up right tooling for debugging production incidents.

System & Application Troubleshooting

- Perform deep troubleshooting for:
- Distributed systems
- Microservices-based architectures on containerised workloads
- Java and Golang applications
- Strong debugging of:
- Application issues
- Infrastructure issues
- Network-related problems

Plan and execute continuous improvement
- Identify and eliminate repetitive manual tasks.
- Drive reliability engineering practices and culture. (DRY – Don’t Repeat Yourself)
- Collaborate with development teams to improve system design and resilience.

Technical Skills

Cloud & Platform

- Strong expertise in Google Cloud Platform (GCP):
- GKE, VPC, IAM, Load Balancing, LB, Certs, KMS, logs and metrics exploration
- BigQuery, Pub/Sub
- Positive understanding of cloud architecture and landing zones

Infrastructure as Code





- Strong hands-on experience with Terraform
- Ability to write and debug Terraform code from scratch

Containers & Orchestration
- Deep expertise in:
- Kubernetes (GKE)
- Docker
- Strong troubleshooting experience in Kubernetes environments

CI/CD & Automation

- Hands-on experience with:
- Jenkins (pipeline-based CI/CD)
- GitHub
- Strong scripting skills:
- Python (preferred)
- Shell scripting
- Experience with automation frameworks and tooling

Observability

- Experience with:
- Dynatrace / Grafana
- Log, metrics, and trace-based monitoring

Programming & Debugging

- Working knowledge of:
- Java and/or Golang applications
- Strong debugging skills across application and infrastructure layers

Linux & Networking

- Strong Linux fundamentals
- Deep understanding of TCP/IP networking
- Ability to debug network issues in distributed systems

Reliability Engineering Skills

- Solid understanding of:
- SLI, SLO, SLA, Error Budgets
- Demonstrable and Proven Experience improving:
- MTTR, MTTA
- Experience handling incident management lifecycle

Soft Skills

- Strong analytical and troubleshooting mindset
- Excellent communication and stakeholder management
- Ability to work in high-pressure production environments
- Ownership-driven and proactive approach

Preferred candidates with:

- Experience in high-scale distributed systems
- Exposure to banking/financial domain (optional but valuable)
- Understanding of security and compliance practices
- Experience with deployment strategies:
- Canary, Blue-Green
- Strong GCP + Kubernetes + Terraform core
- Hands-on production troubleshooting expert
- Good at automation + reducing toil
- Deep understanding of SRE principles
- Comfortable in 24x7 production environments

📌 Site Reliability Engineer (SRE) – GCP Platform (Bengaluru)
🏢 ITC Infotech
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (sre) – gcp platform (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (sre) – gcp platform (bengaluru) / bengaluru