Sr Kubernetes Engineer ( SME ) (Mumbai)

Sr Kubernetes Engineer ( SME ) (Mumbai)

22 Aug
|
Neysa Networks
|
Mumbai

22 Aug

Neysa Networks

Mumbai

Job Description

About the Role:

We are building and running mission-critical production infrastructure on Kubernetes. As a Senior Kubernetes Engineer, you will own the full stack from the underlying OS and container runtime through networking, storage, and the cluster control plane itself. You will engage across the full lifecycle: architecture, deployment, hardening, and day-2 operations for multi-cluster, multi-tenant environments. This is hands-on infrastructure work with direct ownership of production reliability.

What you will be doing:

Design, deploy, and operate production-grade Kubernetes clusters across bare-metal, cloud, and hybrid environments.

Manage cluster lifecycle end-to-end: provisioning, version upgrades, patching, scaling, and capacity planning

Architect and run multi-cluster/multi-tenant setups using Kamaji, Rancher (hosted control planes), and vCluster (virtual clusters)

Configure and troubleshoot CNI plugins (Calico, Cilium) pod networking, network policies, BGP, and eBPF dataplanes.

Own cluster DNS (CoreDNS) configuration, service discovery, and resolution troubleshooting.

Manage container runtime (containerd, CRI-O) and underlying OS: Linux tuning, kernel/sysctl parameters, systemd, cgroups

Deploy and operate service mesh (Istio, Cilium mesh) for traffic management

Manage persistent storage: PV/PVC, StorageClasses, CSI drivers (Rook/Ceph, Longhorn, cloud-native CSI)

Own the network stack: ingress controllers, load balancing, MetalLB/BGP, firewalling, and network troubleshooting

Implement GitOps and Infrastructure-as-Code (ArgoCD/FluxCD, Terraform, Helm) for cluster and workload delivery

Harden clusters: RBAC, Pod Security Standards, network policies, secrets management (Vault, Sealed Secrets), image scanning

Own backup and disaster recovery (Velero, etcd snapshotting) and run DR drills

Provide on-call production support:



monitor cluster health, troubleshoot incidents, and drive root-cause resolution

Create comprehensive documentation, runbooks, and knowledge bases for operational continuity and knowledge transfer

What we need to see:

Core Kubernetes Infrastructure (5+ years)

Deep expertise in Kubernetes architecture: control plane, etcd, kube-apiserver, scheduler, controller-manager, kubelet

Proven experience designing, deploying, and troubleshooting production clusters at scale

Hands-on with multi-cluster/multi-tenant tooling Rancher, Kamaji, and/or vCluster (strongly preferred)

CNI expertise: Calico, Cilium network policy design, BGP, VXLAN/IPIP encapsulation, eBPF

Container runtime internals: containerd, CRI-O, runc configuration and troubleshooting

Robust Linux systems administration: kernel tuning, systemd, cgroups/namespaces, sysctl, package/OS lifecycle management

Storage: PV/PVC, StorageClasses, CSI drivers, Ceph/Rook, Longhorn, NFS

Service mesh experience: Istio, Linkerd, or Cilium service mesh

CoreDNS configuration, custom resolvers, and DNS troubleshooting in cluster environments

Networking depth: ingress controllers (NGINX, Envoy, Traefik), load balancing, MetalLB, and diagnostic tooling (tcpdump, iptables/nftables, conntrack)

Automation Infrastructure-as-Code:

Helm chart authoring and lifecycle management

GitOps workflows: ArgoCD or FluxCD

IaC and configuration management: Terraform, Ansible

CI/CD pipeline integration for cluster and application delivery





Scripting proficiency: Bash and Python (Go a plus)

Ways to stand out from the rest:

Kubernetes certifications: CKA, CKAD, CKS

Production experience with Rancher, Kamaji, and vCluster together (fleet/hosted-control-plane management)

Multi-cloud Kubernetes: EKS, AKS, GKE, and bare-metal

Experience with GPU-enabled clusters and AI/ML workload scheduling

Familiarity with AI/LLM serving stacks on Kubernetes vLLM, KServe, Triton Inference Server, GPU operator/device plugin, MIG partitioning

General AI infrastructure knowledge: model serving patterns, inference autoscaling, GPU scheduling constraints

Contributions to CNCF projects or an active open-source/GitHub presence

Minimum Qualifications:

Bachelor s degree in computer science, Electrical/Computer Engineering, or related field (or equivalent industry experience)

5+ years of hands-on production Kubernetes experience

Demonstrated ownership of CNI, DNS, storage, and networking within Kubernetes environments

Solid grounding in containerd/OS-level troubleshooting

Production on-call and incident-response experience.

Soft Skills:

Strong problem-solving and debugging abilities under production pressure

Ownership mindset with accountability for platform reliability and stability

Cross-functional collaboration with application, security, operations and platform teams

Proactive approach to continuous learning and staying current with the Kubernetes/CNCF ecosystem

Strong documentation and communication skills, with ability to defend design decisions in peer reviews

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Sr Kubernetes Engineer ( SME ) (Mumbai)
🏢 Neysa Networks
📍 Mumbai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: sr kubernetes engineer ( sme ) (mumbai) / mumbai