Senior Kubernetes Engineer (SME) (Mumbai)

Senior Kubernetes Engineer (SME) (Mumbai)

22 Aug
|
Neysa
|
Mumbai

22 Aug

Neysa

Mumbai

About the Role:

We are building and running mission-critical production infrastructure on Kubernetes. As a Senior Kubernetes Engineer, you will own the full stack - from the underlying OS and container runtime through networking, storage, and the cluster control plane itself. You will engage across the full lifecycle: architecture, deployment, hardening, and day-2 operations for multi-cluster, multi-tenant environments. This is hands-on infrastructure work with direct ownership of production reliability.

What you will be doing:



Design, deploy, and operate production-grade Kubernetes clusters across bare-metal, cloud, and hybrid environments.



Manage cluster lifecycle end-to-end: provisioning, version upgrades, patching, scaling, and capacity planning



Architect and run multi-cluster/multi-tenant setups using Kamaji, Rancher (hosted control planes), and vCluster (virtual clusters)



Configure and troubleshoot CNI plugins (Calico, Cilium) — pod networking, network policies, BGP, and eBPF dataplanes.



Own cluster DNS (CoreDNS) configuration, service discovery, and resolution troubleshooting.



Manage container runtime (containerd, CRI-O) and underlying OS: Linux tuning, kernel/sysctl parameters, systemd, cgroups



Deploy and operate service mesh (Istio, Cilium mesh) for traffic management



Manage persistent storage: PV/PVC, StorageClasses, CSI drivers (Rook/Ceph, Longhorn, cloud-native CSI)



Own the network stack: ingress controllers, load balancing, MetalLB/BGP, firewalling, and network troubleshooting



Implement GitOps and Infrastructure-as-Code (ArgoCD/FluxCD, Terraform, Helm) for cluster and workload delivery



Harden clusters: RBAC, Pod Security Standards, network policies, secrets management (Vault, Sealed Secrets), image scanning



Own backup and disaster recovery (Velero, etcd snapshotting) and run DR drills







Provide on-call production support: monitor cluster health, troubleshoot incidents, and drive root-cause resolution



Create comprehensive documentation, runbooks, and knowledge bases for operational continuity and knowledge transfer

What we need to see:



Core Kubernetes & Infrastructure (5+ years)



Deep expertise in Kubernetes architecture: control plane, etcd, kube-apiserver, scheduler, controller-manager, kubelet



Proven experience designing, deploying, and troubleshooting production clusters at scale



Hands-on with multi-cluster/multi-tenant tooling — Rancher, Kamaji, and/or vCluster (strongly preferred)



CNI expertise: Calico, Cilium — network policy design, BGP, VXLAN/IPIP encapsulation, eBPF



Container runtime internals: containerd, CRI-O, runc — configuration and troubleshooting



Strong Linux systems administration: kernel tuning, systemd, cgroups/namespaces, sysctl, package/OS lifecycle management



Storage: PV/PVC, StorageClasses, CSI drivers, Ceph/Rook, Longhorn, NFS



Service mesh experience: Istio, Linkerd, or Cilium service mesh



CoreDNS configuration, custom resolvers, and DNS troubleshooting in cluster environments



Networking depth: ingress controllers (NGINX, Envoy, Traefik), load balancing, MetalLB, and diagnostic tooling (tcpdump, iptables/nftables, conntrack)

Automation & Infrastructure-as-Code:



Helm chart authoring and lifecycle management



GitOps workflows:



ArgoCD or FluxCD



IaC and configuration management: Terraform, Ansible



CI/CD pipeline integration for cluster and application delivery



Scripting proficiency: Bash and Python (Go a plus)

Ways to stand out from the rest:



Kubernetes certifications: CKA, CKAD, CKS



Production experience with Rancher, Kamaji, and vCluster together (fleet/hosted-control-plane management)



Multi-cloud Kubernetes: EKS, AKS, GKE, and bare-metal



Experience with GPU-enabled clusters and AI/ML workload scheduling '



Familiarity with AI/LLM serving stacks on Kubernetes — vLLM, KServe, Triton Inference Server, GPU operator/device plugin, MIG partitioning



General AI infrastructure knowledge: model serving patterns, inference autoscaling, GPU scheduling constraints



Contributions to CNCF projects or an active open-source/GitHub presence

Minimum Qualifications:



Bachelor’s degree in computer science, Electrical/Computer Engineering, or related field (or equivalent industry experience)



5+ years of hands-on production Kubernetes experience



Demonstrated ownership of CNI, DNS, storage, and networking within Kubernetes environments



Solid grounding in containerd/OS-level troubleshooting



Production on-call and incident-response experience.

Soft Skills:



Strong problem-solving and debugging abilities under production pressure



Ownership mindset with accountability for platform reliability and stability



Cross-functional collaboration with application, security, operations and platform teams



Proactive approach to continuous learning and staying current with the Kubernetes/CNCF ecosystem



Solid documentation and communication skills, with ability to defend design decisions in peer reviews

📌 Senior Kubernetes Engineer (SME) (Mumbai)
🏢 Neysa
📍 Mumbai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior kubernetes engineer (sme) (mumbai) / mumbai

Subscribe to this job alert:

Get the latest job offers by email for: senior kubernetes engineer (sme) (mumbai) / mumbai