We are seeking an experienced Cloud Platform Engineer / Site Reliability Engineer (SRE) to support the design, delivery, automation, and operational reliability of cloud platforms within a large-scale banking ecosystem. The role combines platform engineering, DevOps, and reliability engineering to build secure, scalable, and resilient cloud-native services while ensuring operational excellence across production settings.
Key Responsibilities
Design, build, automate, and operate cloud platform capabilities across Google Cloud Platform (GCP) and Kubernetes (GKE) environments.
Manage platform services through their full lifecycle, including implementation, deployment, production support, and continuous improvement.
Drive adoption of Site Reliability Engineering (SRE) practices, focusing on automation, reliability, observability, and reduction of operational toil.
Implement and maintain Infrastructure as Code (IaC) solutions using Terraform and contemporary CI/CD pipelines.
Ensure platform security, scalability, resilience, and compliance with enterprise standards.
Support incident management, problem management, capacity planning, disaster recovery, and production readiness activities.
Collaborate with engineering, product, security,
and platform teams to deliver highly available and supportable services.
Enhance platform observability using monitoring, logging, tracing, and alerting solutions.
Troubleshoot complex issues spanning infrastructure, networking, Kubernetes, cloud services, and integrations.
Required Skills & Experience
Strong experience as a Cloud Platform Engineer, Site Reliability Engineer, Platform Engineer, DevOps Engineer, or Infrastructure Engineer.
Hands-on experience with Google Cloud Platform (GCP) and cloud-native architectures.
Solid expertise in Kubernetes/GKE, including deployments, networking, ingress, scaling, upgrades, and operational support.
Proven experience with Terraform and Infrastructure as Code practices.
Experience building and managing CI/CD pipelines and automation frameworks.
Solid understanding of SRE concepts including SLIs, SLOs, Error Budgets, Reliability Engineering, and Operational Readiness.
Strong experience in incident response, monitoring, alerting, resilience engineering, and service improvement.
Hands-on troubleshooting across cloud infrastructure, networking, security, and platform services.
📌 Devops /sre/cloud Platform Engineer With Gcp Opening With Capco Bengaluru (India)
🏢 CAPCO
📍 India