17 Sep
|
Tata Consultancy Services
|
Greater Noida
17 Sep
Tata Consultancy Services
Greater Noida
Role & responsibilities
We are looking for a highly skilled Site Reliability Engineer (SRE) to build, operate, and continuously improve highly available, scalable, and observable platforms running on bare****metal Kubernetes clusters and Google Kubernetes Engine (GKE).
The ideal candidate brings deep Kubernetes expertise, strong cloud-native experience on GCP, and a passion for reliability, automation, and operational excellence. This role works closely with application, platform, and architecture teams to ensure production systems are resilient, secure, and performant at scale.
Key Responsibilities
- Design, operate, and support Kubernetes platforms across bare****metal clusters and GKE
- Ensure high availability, scalability, performance, and reliability of production systems
- Implement and manage GitOps-based deployment workflows using tools like Argo CD
- Build, maintain, and optimize CI/CD pipelines using tools such as GitHub Actions, Harness, CircleCI, or equivalent
- Deploy and manage applications using Helm, including canary and progressive delivery strategies
- Implement comprehensive observability using Prometheus, Grafana, Loki, and Tempo
- Proactively monitor systems, troubleshoot incidents, and perform root cause analysis (RCA)
- Partner with development teams to improve service reliability, scalability, and operational maturity
- Provision and manage cloud infrastructure on Google Cloud Platform (GCP)
- Automate infrastructure and platform operations using Infrastructure as Code (IaC) and scripting
- Drive continuous improvements in resilience, automation, and operational efficiency
Required Skills & Qualifications
- Strong hands-on experience with Kubernetes architecture and administration
- Experience managing both bare-metal Kubernetes clusters and Google Kubernetes Engine (GKE)
- Solid understanding of Google Cloud Platform (GCP) services and networking concepts
- Proven experience with GitOps practices and tools such as Argo CD
- Proficiency with CI/CD tools (GitHub Actions, Harness, CircleCI, or similar)
- Practical experience with:
- Helm
- Canary / progressive deployments
- Strong expertise in observability and monitoring:
- Prometheus
- Grafana
- Loki
- Tempo
- Experience with Terraform for infrastructure provisioning
- Understanding of modern API technologies such as GraphQL
- Familiarity with API management platforms (Apigee Edge, Apigee X)
- Knowledge of CDN and edge services (e.g., Akamai)
Valuable to Have
- Working knowledge of Java (Spring Boot) and/or Node.js framework
- Understanding of microservices architecture and service-to-service communication
- Experience with Ansible or similar configuration management tools
- Exposure to hybrid or multi****cloud environments
- Experience in performance tuning and cost optimization on GCP
- Understanding of Kubernetes and cloud security best practices
- SRE experience aligned with SLIs, SLOs, and error budgets
📌 Walk-in || Site Reliability Engineer (Greater Noida)
🏢 Tata Consultancy Services
📍 Greater Noida