TCS hiring for GCP - Site Reliability Engineer (SRE)
Location - Bangalore / Chennai
Experience - 8 - 12 years
Notice Period - 0 - 30 days
Key Responsibilities -
Reliability & Operations
Own end-to-end production systems reliability, availability, scalability, cost and performance.
Drive measurable improvements in MTTR, MTTA, and incident response practices using automation and runbook additions and process enhancements.
Participate in 24x7 on-call rotations and handle high-severity incidents and document the learnings on ongoing basis.
Establish and manage SLI, SLO, SLA, Error Budgets, and operational metrics for mission critical services and partner with engineering teams with full accountability for upholding the SLOs.
Partner with the various engineering, operations and cloud management teams to deliver highly reliable service in a timely manner.
Cloud & Infrastructure
Design, deploy, and manage infrastructure on Google Cloud Platform (GCP).
Work extensively on:
GKE (Kubernetes Engine)
Compute, networking, IAM, Load Balancers, TLS Certs
BigQuery, Pub/Sub, cloud logging enhancement, metrics and logs analysis
Implement and manage infrastructure using Terraform (Infrastructure as Code).