TCS hiring for GCP - Site Reliability Engineer (SRE)
Location - Bangalore / Chennai
Experience - 8 - 12 years
Notice Period - 0 - 30 days
Key Responsibilities -
Reliability & Operations
- Own end-to-end production systems reliability, availability, scalability, cost and performance.
- Drive measurable improvements in MTTR, MTTA, and incident response practices using automation and runbook additions and process enhancements.
- Participate in 24x7 on-call rotations and handle high-severity incidents and document the learnings on ongoing basis.
- Establish and manage SLI, SLO, SLA, Error Budgets, and operational metrics for mission critical services and partner with engineering teams with full accountability for upholding the SLOs.
- Partner with the various engineering, operations and cloud management teams to deliver highly reliable service in a timely manner.
Cloud & Infrastructure
- Design, deploy, and manage infrastructure on Google Cloud Platform (GCP).
- Work extensively on:
- GKE (Kubernetes Engine)
- Compute, networking, IAM, Load Balancers, TLS Certs
- BigQuery, Pub/Sub, cloud logging enhancement, metrics and logs analysis
- Implement and manage infrastructure using Terraform (Infrastructure as Code).