Senior Kubernetes Platform Engineer (Bengaluru)

Senior Kubernetes Platform Engineer (Bengaluru)

21 Aug
|
Aivar Innovations
|
Bengaluru

21 Aug

Aivar Innovations

Bengaluru

About Us

Aivar Innovations is an AI-native services company and AWS Preferred Partner building governed, production-grade agentic AI systems. We partner with enterprises to deploy intelligent agents that automate complex business processes—from intelligent customer interactions to enterprise knowledge systems.

Experience: 5–9 years | 4+ years building or operating production Kubernetes platforms, controllers, operators, or cloud-native infrastructure The Role: You build the Kubernetes substrate that makes Kubogent possible.

Kubogent is a Kubernetes-native AI infrastructure and MLOps platform. A central control plane manages multiple workload clusters, allocates infrastructure to tenants and projects, deploys platform capabilities on demand, runs GPU and ML workloads, exposes remote operational access, and continuously reconciles desired state with what is actually running.

This role is for someone who understands Kubernetes as a distributed system and programmable control plane, not just as a deployment target.

What You'll Do

- Own Kubernetes control loops. Design and build CRDs, controllers, operators, reconcilers, finalizers, watches, status models, and lifecycle state machines using Go and controller-runtime.
- Build multi-cluster control. Help design and implement how Kubogent onboards, authenticates, observes, upgrades, and controls workload clusters that may only initiate outbound connections to the control plane.
- Build the workload-cluster agent. Own registration, heartbeat, inventory, command execution, reconnect behaviour, versioning, rollout, credential rotation, and failure recovery.

- Translate platform intent into Kubernetes state. Convert concepts such as project placement, capabilities, resource allocation, model deployments, notebooks, training jobs, and shared services into protected, idempotent Kubernetes reconciliation.
- Own resource isolation and placement. Work with namespaces, quotas, limits, priority, scheduling, taints/tolerations, affinity, topology, GPU resources, gang scheduling, and workload placement.
- Build GPU and accelerator support.

Integrate with NVIDIA GPU Operator, device plugins, MIG where appropriate, node feature discovery, topology-aware scheduling, and accelerator-specific runtime requirements.

- Build secure remote operations.



Design mechanisms for browser-based kubectl/exec/log access, tunneled cluster connectivity, least-privilege credentials, mTLS, authorization, and auditable command execution.
- Own cluster capability deployment. Build mechanisms that install only the operators and services required by capabilities enabled on projects or clusters, rather than treating every cluster as identical.
- Design for unreliable environments.

Clusters disappear, links break, agents restart, watches expire, APIs throttle, upgrades partially fail, and reconciliation gets repeated. Your systems must remain correct anyway.

- Work deeply with Kubernetes API machinery. Informers, watches, admission, status conditions, server-side apply, resource versions, optimistic concurrency, garbage collection, RBAC, and API conventions.
- Own platform upgrades. Design safe version skew, agent upgrades, CRD evolution, migration, backward
- compatibility, and rollout/rollback strategies.
- Drive observability for the platform itself.

Instrument operators and agents with logs, metrics, traces, health checks, queue depth, reconciliation latency, and actionable failure signals.

- Set the bar. Review designs and code, mentor engineers, and establish patterns for building reliable Kubernetes-native systems.

What We're Looking For

- 5+ years in software, platform, infrastructure, SRE, or cloud engineering, with substantial hands-on Kubernetes experience.
- Strong Go. You should be comfortable designing production services, concurrency, interfaces, testing, profiling, and failure handling in Go.
- Deep Kubernetes internals. Controllers, CRDs, reconciliation, API machinery, watches/informers, RBAC, admission, scheduling, storage, networking, and workload lifecycle.
- You have built Kubernetes software, not only operated clusters.

We especially value experience with Kubebuilder, controller-runtime, Operator SDK, custom schedulers,



admission webhooks, or Kubernetes-integrated platforms.

- Strong distributed-systems instincts. Idempotency, retries, at-least-once execution, eventual consistency, leases, leader election, partial failure, backpressure, and state convergence should be familiar ideas.
- Experience operating Kubernetes across environments. Cloud-managed Kubernetes and on-prem/private-cloud experience are both valuable.
- Comfort with networking.

TCP/TLS, mTLS, proxies, reverse tunnels, WebSockets or streaming RPC, DNS, load balancers, ingress/gateway, and debugging connectivity failures.

- Production troubleshooting ability. You can move from symptom to root cause across controllers, API servers, networking, scheduling, container runtimes, storage, and workloads.
- Strong security fundamentals. Service identities, certificates, RBAC, secrets, credential rotation, least privilege, tenant isolation, and auditability.
- Fluency with agentic coding tools.

We expect AI to accelerate implementation and investigation while you remain responsible for architecture, correctness, failure handling, and operational quality.

Strong Pluses

- Multi-cluster management platforms.
- Kubernetes API aggregation or extension patterns.
- Cluster API, Crossplane, Argo CD, Flux, Rancher, Rafay, Open Cluster Management, or similar systems.
- Envoy, reverse tunnels, relay systems, or secure remote cluster access.
- GPU scheduling, NVIDIA GPU Operator, MIG, CUDA-aware workloads, or distributed training infrastru
- Prometheus, OpenTelemetry, Loki, or Kubernetes observability stacks.
- Kubernetes conformance, upgrade testing, chaos testing, or large-scale fleet management.

This Role May Not Be The Best Match If

- Your Kubernetes experience is primarily writing manifests, Helm charts, and maintaining CI/CD pipelines. We arebuilding Kubernetes-native control-plane software.
- You expect infrastructure to be reliable and synchronous. Disconnected clusters, duplicate commands, partial
- upgrades, stale state, and repeated reconciliation are normal operating conditions here.
- You prefer solving platform problems by adding manual operational procedures. Kubogent must turn those procedures into productized, automated control loops.

📌 Senior Kubernetes Platform Engineer (Bengaluru)
🏢 Aivar Innovations
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior kubernetes platform engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: senior kubernetes platform engineer (bengaluru) / bengaluru