01 Oct
|
CAPCO
|
Bengaluru
We are seeking an experienced Cloud Platform Engineer / Site Reliability Engineer (SRE) to support the design, delivery, automation, and operational reliability of cloud platforms within a large-scale banking ecosystem. The role combines platform engineering, DevOps, and reliability engineering to build secure, scalable, and resilient cloud-native services while ensuring operational excellence across production environments.
Key Responsibilities
- Design, build, automate, and operate cloud platform capabilities across Google Cloud Platform (GCP) and Kubernetes (GKE) environments.
- Manage platform services through their full lifecycle, including implementation, deployment, production support, and continuous improvement.
- Drive adoption of Site Reliability Engineering (SRE) practices, focusing on automation, reliability, observability, and reduction of operational toil.
- Implement and maintain Infrastructure as Code (IaC) solutions using Terraform and modern CI/CD pipelines.
- Ensure platform security, scalability, resilience, and compliance with enterprise standards.
- Support incident management, problem management, capacity planning, disaster recovery, and production readiness activities.
- Collaborate with engineering, product, security,
and platform teams to deliver highly available and supportable services.
- Enhance platform observability using monitoring, logging, tracing, and alerting solutions.
- Troubleshoot complex issues spanning infrastructure, networking, Kubernetes, cloud services, and integrations.
Required Skills & Experience
- Strong experience as a Cloud Platform Engineer, Site Reliability Engineer, Platform Engineer, DevOps Engineer, or Infrastructure Engineer.
- Hands-on experience with Google Cloud Platform (GCP) and cloud-native architectures.
- Solid expertise in Kubernetes/GKE, including deployments, networking, ingress, scaling, upgrades, and operational support.
- Proven experience with Terraform and Infrastructure as Code practices.
- Experience building and managing CI/CD pipelines and automation frameworks.
- Solid understanding of SRE concepts including SLIs, SLOs, Error Budgets, Reliability Engineering, and Operational Readiness.
- Strong experience in incident response, monitoring, alerting, resilience engineering, and service improvement.
- Hands-on troubleshooting across cloud infrastructure, networking, security, and platform services.
📌 Devops /SRE/Cloud Platform Engineer with GCP Opening with CAPCO (Bengaluru)
🏢 CAPCO
📍 Bengaluru