Platform Engineer (Bengaluru)

Platform Engineer (Bengaluru)

01 Oct
|
Dhruv Technology Solutions
|
Bengaluru

01 Oct

Dhruv Technology Solutions

Bengaluru

Core Responsibility Summary

Deploys, operates, and evolves the Kafka / Flink / ClickHouse platform on Kubernetes across multiple factory data centers. Ensures the platform is reliable, observable, performant, and ready for data engineering teams to build on top of.

Kubernetes & Container Platform

- Deep Kubernetes expertise (workload types, scheduling, resource requests/limits, namespaces, RBAC)
- Kubernetes operators for stateful systems — Strimzi (Kafka), Flink Kubernetes Operator, ClickHouse Operator (Altinity)
- Helm chart authoring and management (not just consumption — writing and maintaining charts)
- StatefulSet management, persistent volume claims, storage class configuration
- Multi-cluster and multi-site Kubernetes deployments
- Node affinity, taints/tolerations, pod disruption budgets for workload isolation
- Kubernetes networking: services, ingress, network policies, DNS
- Secrets management (Kubernetes secrets, Vault integration, or equivalent)

Platform Component Expertise

Kafka

- Kafka cluster sizing, broker configuration, topic design (partitions, replication, retention)
- Strimzi operator — KafkaNodePool, KafkaTopic, KafkaUser, KafkaMirrorMaker2 CRDs
- Kafka Connect cluster deployment and connector lifecycle management
- Kafka security: TLS, mTLS, SASL, ACLs
- Kafka upgrade and rolling restart procedures with zero data loss

Flink

- Flink cluster deployment modes on Kubernetes (Application mode, Session mode)
- Flink Kubernetes Operator — FlinkDeployment, FlinkSessionJob CRDs
- Checkpointing and savepoint management (S3 or distributed storage backends)
- Flink HA configuration (Kubernetes-native HA or ZooKeeper)
- Job submission, monitoring, and restarts without data loss

ClickHouse

- ClickHouse cluster deployment with Altinity operator or ClickHouse Operator
- Shard and replica topology design for multi-factory deployments
- Storage configuration: tiered storage,



disk policies, S3-backed cold storage
- Backup and restore strategies
- ClickHouse Keeper vs ZooKeeper for coordination

Observability & Monitoring

- Full observability stack deployment: Prometheus, Grafana, Alertmanager
- Kafka metrics: broker JMX metrics, consumer lag (Burrow or Cruise Control), under-replicated partitions
- Flink metrics: checkpoint duration, backpressure, restart frequency, throughput/latency per operator
- ClickHouse metrics: query latency, merge backlog, insert rate, replication lag, memory/disk pressure
- Kubernetes cluster metrics: node utilization, pod restarts, OOMKill events, PVC utilization
- SLO/SLI definition and alerting for platform uptime and data freshness
- Distributed tracing for pipeline end-to-end latency visibility
- Log aggregation (EFK stack, Loki, or equivalent) for platform component logs

Load Testing & Capacity Planning

- Load testing frameworks for streaming platforms (Gatling, k6, custom Kafka producer harnesses)
- Designing load test scenarios that simulate factory production volumes and burst conditions
- Benchmarking Kafka throughput (MB/s, msg/s) under varying partition counts and consumer configurations
- Flink job stress testing — measuring backpressure onset, checkpoint interval degradation under load
- ClickHouse ingestion benchmarking (insert throughput, concurrent query performance under load)
- Capacity modeling: translating factory data volumes and growth projections into infrastructure sizing




- Breaking point analysis — identifying bottlenecks before production traffic hits them

Isolation & Multi-Tenancy

- Namespace-level isolation for multiple factory deployments on shared Kubernetes clusters
- Kafka topic and consumer group isolation strategies across factory sites and teams
- Resource quotas and LimitRanges to prevent noisy-neighbor problems between factory workloads
- Network policy enforcement to isolate factory data plane traffic
- Separate Kafka clusters vs shared clusters with namespace isolation — trade-off analysis and implementation
- Per-factory ClickHouse cluster deployment vs shared cluster with database/user isolation

CI/CD & GitOps for Platform

- GitOps-driven platform deployment: ArgoCD or Flux for Kubernetes manifest management
- Helm chart versioning and promotion across environments (dev → staging → factory-prod)
- Operator upgrade management and CRD migration procedures
- Automated smoke tests post-deployment to validate platform health
- Change management for platform updates in production factory environments

Security & Compliance

- mTLS and certificate management (cert-manager on Kubernetes)
- Kafka ACL management at scale
- Network segmentation between factory OT and IT data flows
- Image scanning and supply chain security for platform container images
- Audit logging for platform access and configuration changes

AI-Assisted Development

- Using AI coding assistants (GitHub Copilot, Cursor, or equivalent) to accelerate Helm chart authoring, operator configuration, and runbook generation
- AI-assisted troubleshooting — using LLMs to diagnose Kubernetes events, Flink exceptions, and Kafka error logs faster
- Generating observability dashboards and alert rules with AI assistance
- Critical review of AI-generated infrastructure code before applying to production

📌 Platform Engineer (Bengaluru)
🏢 Dhruv Technology Solutions
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: platform engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: platform engineer (bengaluru) / bengaluru