19 Aug
|
Neurealm
|
Chennai
Key Responsibilities
Cluster Operations & Support
- Manage and support Kubernetes clusters (on Azure AKS) including day-to-day operations, incident and problem management (L2 escalations from L1), and RCA within defined SLAs.
- Perform cluster upgrades (control plane and worker nodes) including version planning, compatibility checks, rollback strategy, and minimizing downtime during upgrades.
- Manage Control Plane components kube-apiserver, kube-scheduler, kube-controller-manager, cloud-controller-manager including health checks, HA configuration, and troubleshooting component failures.
- Monitor, backup, restore, and troubleshoot etcd (cluster state store), including snapshotting, defragmentation, quorum issues, and disaster recovery.
- Perform hands-on operations using kubectl get/describe/logs/exec/apply/patch/rollout/drain/cordon/top/debug for day-to-day administration and troubleshooting.
Troubleshooting
- Troubleshoot Pod-level issues: CrashLoopBackOff, ImagePullBackOff, OOMKilled, pending pods, failed scheduling, readiness/liveness probe failures, and node pressure conditions.
- Diagnose and resolve Networking issues: DNS resolution (CoreDNS), service discovery, pod-to-pod communication, service ClusterIP/NodePort/LoadBalancer issues, and network policy conflicts.
- Debug CNI (Container Network Interface) issues — Calico, Cilium, Flannel, Azure CNI, or AWS VPC CNI — including IP exhaustion, plugin failures, and overlay/underlay network issues.
- Troubleshoot Storage issues: PV/PVC binding failures, StorageClass misconfigurations, CSI driver issues, volume mount errors, and data persistence problems across StatefulSets.
- Resolve Ingress issues: Ingress controller (NGINX, Traefik, HAProxy, Azure App Gateway) misconfigurations, TLS/certificate issues, routing rules,
and load balancing problems.
- Investigate node-level issues (disk pressure, memory pressure, kubelet health) and cluster autoscaler behavior.
Security & Access
- Implement and manage RBAC (Roles, ClusterRoles, RoleBindings, ClusterRoleBindings) and Service Accounts for secure, least-privilege access.
- Manage Network Policies for pod-to-pod traffic control and micro-segmentation.
- Support Pod Security Standards/Admission, Secrets management, and integration with external secret stores (Vault, Azure Key Vault, AWS Secrets Manager).
- Ensure compliance with security best practices, image scanning, and Kubernetes CIS benchmark adherence.
Platform & Automation
- Contribute to Platform Design decisions — cluster architecture, multi-tenancy models, namespace strategy, resource quotas/limits, and capacity planning.
- Design and maintain Helm charts, Kustomize overlays, and GitOps workflows (ArgoCD/Flux) for application and platform deployments.
- Work with Infrastructure as Code (Terraform, Bicep, CloudFormation) for provisioning clusters and supporting cloud resources.
- Implement and maintain CI/CD pipelines integrating with Kubernetes deployments (Azure DevOps, GitHub Actions, Jenkins, GitLab CI).
- Configure and maintain observability stack — Prometheus, Grafana, Loki/ELK, Azure Monitor/Container Insights — for cluster and workload monitoring.
- Manage service mesh (Istio/Linkerd) where applicable for traffic management, mTLS, and observability.
- Document runbooks, SOPs, architecture diagrams, and knowledge base articles for recurring issues.
- Participate in on-call rotation and support during planned maintenance/outages.
Required Skills & Qualifications
- Bachelor's degree in computer science, IT, or related field (or equivalent experience).
- 2-3 years of hands-on experience in Kubernetes administration/engineering.
- Strong command over kubectl and Kubernetes object model (Deployments, StatefulSets, DaemonSets, Jobs/CronJobs, ConfigMaps, Secrets).
- Deep troubleshooting experience across Pods, Networking, Storage, and Ingress.
- Solid understanding of Control Plane architecture and etcd operations/backup-restore.
Preferred / Good to Have
- Certified Kubernetes Administrator (CKA) or Certified Kubernetes Security Specialist (CKS).
- Experience with Platform Design for multi-cluster/multi-tenant environments.
- Exposure to service mesh (Istio, Linkerd) and API gateways.
- Experience with managed Kubernetes services (AKS, EKS, GKE) alongside on-prem/bare-metal clusters.
- Knowledge of cost optimization tools (Kubecost) and cluster autoscaling strategies.
- Experience with chaos engineering / resilience testing tools.
Soft Skills
- Strong communication skills to collaborate with Dev, SRE, Security, and Network teams.
- Ability to work independently under SLA pressure in a 24x7 support setting.
- Analytical, methodical troubleshooting approach with an ownership mindset.
- Continuous learning attitude given the fast-evolving Kubernetes ecosystem.
Note: "Preferred Immediate to 15 days Joiners"
📌 Azure Kubernetes Engineer (Chennai)
🏢 Neurealm
📍 Chennai