About the Role:
We are building and running mission-critical production infrastructure on Kubernetes. As a Senior Kubernetes Engineer, you will own the full stack - from the underlying OS and container runtime through networking, storage, and the cluster control plane itself. You will engage across the full lifecycle: architecture, deployment, hardening, and day-2 operations for multi-cluster, multi-tenant environments. This is hands-on infrastructure work with direct ownership of production reliability.
What you will be doing:
- Design, deploy, and operate production-grade Kubernetes clusters across bare-metal, cloud, and hybrid environments.
- Manage cluster lifecycle end-to-end: provisioning, version upgrades, patching, scaling, and capacity planning
- Architect and run multi-cluster/multi-tenant setups using Kamaji, Rancher (hosted control planes), and vCluster (virtual clusters)
- Configure and troubleshoot CNI plugins (Calico, Cilium) — pod networking, network policies, BGP, and eBPF dataplanes.
- Own cluster DNS (CoreDNS) configuration, service discovery, and resolution troubleshooting.
- Manage container runtime (containerd, CRI-O) and underlying OS: Linux tuning, kernel/sysctl parameters, systemd, cgroups
- Deploy and operate service mesh (Istio, Cilium mesh) for traffic management
- Manage persistent storage: PV/PVC, StorageClasses, CSI drivers (Rook/Ceph, Longhorn, cloud-native CSI)
- Own the network stack: ingress controllers, load balancing, MetalLB/BGP, firewalling, and network troubleshooting
- Implement GitOps and Infrastructure-as-Code (ArgoCD/FluxCD, Terraform, Helm) for cluster and workload delivery
- Harden clusters: RBAC, Pod Security Standards, network policies, secrets management (Vault, Sealed Secrets), image scanning
- Own backup and disaster recovery (Velero, etcd snapshotting)
and run DR drills
- Provide on-call production support: monitor cluster health, troubleshoot incidents, and drive root-cause resolution
- Create comprehensive documentation, runbooks, and knowledge bases for operational continuity and knowledge transfer
What we need to see:
- Core Kubernetes & Infrastructure (5+ years)
- Deep expertise in Kubernetes architecture: control plane, etcd, kube-apiserver, scheduler, controller-manager, kubelet
- Proven experience designing, deploying, and troubleshooting production clusters at scale
- Hands-on with multi-cluster/multi-tenant tooling — Rancher, Kamaji, and/or vCluster (strongly preferred)
- CNI expertise: Calico, Cilium — network policy design, BGP, VXLAN/IPIP encapsulation, eBPF
- Container runtime internals: containerd, CRI-O, runc — configuration and troubleshooting
- Strong Linux systems administration: kernel tuning, systemd, cgroups/namespaces, sysctl, package/OS lifecycle management
- Storage: PV/PVC, StorageClasses, CSI drivers, Ceph/Rook, Longhorn, NFS
- Service mesh experience: Istio, Linkerd, or Cilium service mesh
- CoreDNS configuration, custom resolvers, and DNS troubleshooting in cluster environments
- Networking depth: ingress controllers (NGINX, Envoy, Traefik), load balancing, MetalLB, and diagnostic tooling (tcpdump, iptables/nftables, conntrack)
Automation & Infrastructure-as-Code:
- Helm chart authoring and lifecycle management
- GitOps workflows: ArgoCD or FluxCD
- IaC and configuration management: Terraform, Ansible
- CI/CD pipeline integration for cluster and application delivery
- Scripting proficiency: Bash and Python (Go a plus)
Ways to stand out from the rest:
- Kubernetes certifications: CKA, CKAD, CKS
- Production experience with Rancher, Kamaji, and vCluster together (fleet/hosted-control-plane management)
- Multi-cloud Kubernetes: EKS, AKS, GKE, and bare-metal
- Experience with GPU-enabled clusters and AI/ML workload scheduling '
- Familiarity with AI/LLM serving stacks on Kubernetes — vLLM, KServe, Triton Inference Server, GPU operator/device plugin, MIG partitioning
- General AI infrastructure knowledge: model serving patterns, inference autoscaling, GPU scheduling constraints
- Contributions to CNCF projects or an active open-source/GitHub presence
Minimum Qualifications:
- Bachelor’s degree in computer science, Electrical/Computer Engineering, or related field (or equivalent industry experience)
- 5+ years of hands-on production Kubernetes experience
- Demonstrated ownership of CNI, DNS, storage, and networking within Kubernetes environments
- Solid grounding in containerd/OS-level troubleshooting
- Production on-call and incident-response experience.
Soft Skills:
- Strong problem-solving and debugging abilities under production pressure
- Ownership mindset with accountability for platform reliability and stability
- Cross-functional collaboration with application, security, operations and platform teams
- Proactive approach to continuous learning and staying current with the Kubernetes/CNCF ecosystem
- Solid documentation and communication skills, with ability to defend design decisions in peer reviews
📌 Senior Kubernetes Engineer (SME) (Mumbai)
🏢 Neysa
📍 Mumbai