GKE DevOps Engineer (Bengaluru)

GKE DevOps Engineer (Bengaluru)

24 Sep
|
Murphi.ai
|
Bengaluru

24 Sep

Murphi.ai

Bengaluru

GKE DevOps Engineer

Experience: 3–5 Years

Location: Bengaluru – Work from Office (HSR Layout)

Employment Type: Full-time

Primary Platform: Google Cloud Platform (GCP) / Google Kubernetes Engine (GKE)

About Murphi

Murphi is an AI-powered enterprise technology platform serving customers across Healthcare and Banking, Financial Services & Insurance (BFSI) .

In Healthcare (US) , Murphi helps organizations automate clinical, operational, revenue-cycle, compliance, patient-engagement, and documentation workflows using AI. In BFSI (India) , Murphi provides secure enterprise communication, workflow automation, systems integration, and AI capabilities for regulated financial institutions.

Across both industries, Murphi is building secure, scalable, cloud-native platforms combining AI, automation, data, and enterprise workflows to modernize mission-critical operations.

About the Role

We are looking for a hands-on GKE DevOps Engineer to build, operate, secure, scale, and continuously improve Murphi's cloud platform on Google Cloud Platform.

You will support business-critical, microservices-based SaaS environments across Development, QA, and Production , with responsibility for GKE, CI/CD, Infrastructure as Code, security, networking, observability, reliability, scalability, and cloud cost optimization .

Our technology environment includes GKE/GKE Autopilot, Docker, Helm, Terraform, GitHub Actions, Google Cloud Build, Artifact Registry, Workload Identity, Secret Manager, Cloud SQL/PostgreSQL, Redis, RabbitMQ, Cloud Storage, OpenTelemetry, and Vertex AI .

This is a highly hands-on engineering and production-operations role. You should be comfortable working directly with production systems, diagnosing complex issues, automating repetitive work, and taking ownership from problem identification through resolution.

US & Global Production Support

Murphi serves global customers, with significant production operations in the United States . This role will support US production systems and work closely with engineering and operations teams across India and US time zones.

Availability to work during evening and, when required, night hours in India is critical for this position. The candidate must be comfortable with planned overlap with US business hours, production releases, maintenance activities, incident response, and critical production issues.

Key Responsibilities

GKE & Kubernetes

- Operate and improve GKE clusters and Kubernetes workloads across Development, QA, and Production.
- Take operational responsibility for infrastructure supporting US and global customers.
- Manage multi-service and multi-namespace Kubernetes environments.
- Configure Deployments, Services, ConfigMaps, Secrets, RBAC, Jobs, Gateway/Ingress resources, health probes, HPA, Pod Disruption Budgets, and resource requests/limits.
- Troubleshoot pod failures, CPU/memory issues, networking, latency, deployment failures, and resource contention.
- Establish safe deployment, rollback, recovery, and change-management practices.

CI/CD & Infrastructure as Code
- Build and maintain CI/CD pipelines using GitHub Actions and Google Cloud Build .
- Automate build, testing, container creation, security scanning,



publishing, deployment, validation, and rollback.
- Manage container images using Google Artifact Registry and immutable image-digest-based deployments.
- Package and deploy applications using Helm .
- Provision and manage GCP infrastructure using Terraform .
- Develop reusable Terraform modules, manage state, review plans, and maintain separation between Development, QA, and Production.
- Eliminate manual infrastructure configuration and configuration drift.

Cloud Security
- Implement secure workload authentication using GKE Workload Identity, Google IAM, Kubernetes service accounts, and GitHub OIDC Workload Identity Federation .
- Apply least-privilege access across workloads, users, pipelines, and GCP services.
- Manage secrets using Google Secret Manager and External Secrets Operator .
- Implement RBAC, NetworkPolicies, private connectivity, audit logging, IAM reviews, vulnerability remediation, and container security.
- Maintain security practices appropriate for regulated Healthcare and BFSI environments.

Networking

Configure and troubleshoot

- VPCs and private connectivity
- GKE Gateway API / HTTP routing
- Google Cloud Load Balancing
- DNS and TLS certificates
- Firewall rules and NetworkPolicies
- Cloud Armor / WAF

Diagnose service-to-service connectivity, DNS, routing, certificate, firewall, gateway, and load-balancer issues. Reliability, Scalability & Performance
- Design and improve the platform for increasing traffic and concurrent workloads.
- Improve resilience through autoscaling, health checks, resource management, graceful failure handling, and disruption policies.
- Participate in capacity planning, load testing, performance testing, and production-readiness reviews .
- Identify infrastructure and application bottlenecks before they become production incidents.
- Develop operational runbooks, recovery procedures, and troubleshooting documentation.

Monitoring & Incident Management
- Monitor systems using Google Cloud Logging and Cloud Monitoring .
- Implement observability using OpenTelemetry, Prometheus, Grafana, and distributed tracing .
- Build actionable dashboards and alerts covering availability, latency, errors, Kubernetes health, CPU/memory utilization, and application performance.
- Participate actively in production incident response, including incidents during US operating hours.
- Conduct/support root-cause analysis (RCA) and corrective actions.
- Help establish appropriate SLIs and SLOs .

Cloud Cost Optimization
- Regularly review GCP costs across GKE, compute, storage, databases, networking, and monitoring.
- Identify unused, underutilized, or over-provisioned infrastructure.
- Optimize Kubernetes resources, autoscaling, storage, log retention, and managed services.
- Balance performance, reliability, security, scalability, and cost .





Application & Data Platform Support Work with engineering teams supporting:

Cloud SQL/PostgreSQL • Redis • RabbitMQ • Cloud Storage • REST APIs • Microservices • AI/ML workloads • Vertex AI

You should be able to systematically determine whether a production issue originates from the application, container, Kubernetes, database, network, CI/CD pipeline, or GCP infrastructure .

Required Skills & Experience

- 3–5 years in DevOps, SRE, Cloud Engineering, or Platform Engineering.
- Hands-on responsibility for production Kubernetes environments .
- Strong practical experience with GCP and GKE .
- Strong Kubernetes knowledge: Deployments, Services, namespaces, ConfigMaps, Secrets, RBAC, Gateway/Ingress, health probes, autoscaling, and troubleshooting.
- Strong hands-on skills with kubectl, Docker, Helm, Kubernetes YAML, and Artifact Registry .
- CI/CD experience with GitHub Actions and/or Google Cloud Build .
- Strong Terraform experience, including modules, state management, plan reviews, and multi-environment configuration.
- Solid understanding of GCP IAM, service accounts, Workload Identity, OIDC federation, Secret Manager, and least-privilege security .
- Knowledge of VPCs, private connectivity, load balancing, DNS, TLS, firewalls, and Cloud Armor/WAF .
- Experience with HPA, resource management, health checks, deployment strategies, scaling, and production troubleshooting.
- Comfortable with Linux and scripting using Bash, Python, PowerShell, or similar.
- Strong troubleshooting, communication, documentation, and change-management skills.
- Ability to collaborate effectively with Development, QA, and engineering teams.
- Willingness and flexibility to work evening/night hours in India when required to support US production systems, releases, global customers, and critical incidents.
- GKE Autopilot and GKE Gateway API
- External Secrets Operator and KEDA
- Managed Prometheus, OpenTelemetry Collector, Grafana, and distributed tracing
- Cloud SQL/PostgreSQL, Redis, RabbitMQ, and Cloud Storage
- GitOps practices
- SLI/SLO implementation, incident management, disaster recovery, and backup/recovery testing
- Container vulnerability scanning and software supply-chain security
- Vertex AI / AI-ML workloads on GCP
- Experience with regulated Healthcare, Financial Services, Insurance, or Banking platforms
- Familiarity with HIPAA, SOC 2, ISO 27001 , or similar frameworks
- Google Cloud certifications

What We Are Looking For We are looking for a hands-on, dependable engineer who takes ownership of production systems .

You should be comfortable troubleshooting live systems under pressure, automating repetitive work, and working across the full stack from application → container → Kubernetes → network → database → GCP infrastructure .

You should understand that DevOps is not simply about deployment—it is about reliability, automation, security, scalability, observability, performance, and cost .

Most importantly, you should have the ownership and operational discipline required to support mission-critical production systems and global customers across India and US time zones .

📌 GKE DevOps Engineer (Bengaluru)
🏢 Murphi.ai
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: gke devops engineer (bengaluru) / bengaluru