01 Oct
|
Dhruv Technology Solutions
|
Bengaluru
01 Oct
Dhruv Technology Solutions
Bengaluru
Core Responsibility Summary
Deploys, operates, and evolves the Kafka / Flink / ClickHouse platform on Kubernetes across multiple factory data centers. Ensures the platform is reliable, observable, performant, and ready for data engineering teams to build on top of.
Kubernetes & Container Platform
- Deep Kubernetes expertise (workload types, scheduling, resource requests/limits, namespaces, RBAC)
- Kubernetes operators for stateful systems — Strimzi (Kafka), Flink Kubernetes Operator, ClickHouse Operator (Altinity)
- Helm chart authoring and management (not just consumption — writing and maintaining charts)
- StatefulSet management, persistent volume claims, storage class configuration
- Multi-cluster and multi-site Kubernetes deployments
- Node affinity, taints/tolerations, pod disruption budgets for workload isolation
- Kubernetes networking: services, ingress, network policies, DNS
- Secrets management (Kubernetes secrets, Vault integration, or equivalent)
Platform Component Expertise
Kafka
- Kafka cluster sizing, broker configuration, topic design (partitions, replication, retention)
- Strimzi operator — KafkaNodePool, KafkaTopic, KafkaUser, KafkaMirrorMaker2 CRDs
- Kafka Connect cluster deployment and connector lifecycle management
- Kafka security: TLS, mTLS, SASL, ACLs
- Kafka upgrade and rolling restart procedures with zero data loss
Flink
- Flink cluster deployment modes on Kubernetes (Application mode, Session mode)
- Flink Kubernetes Operator — FlinkDeployment, FlinkSessionJob CRDs
- Checkpointing and savepoint management (S3 or distributed storage backends)
- Flink HA configuration (Kubernetes-native HA or ZooKeeper)
- Job submission, monitoring, and restarts without data loss
ClickHouse
- ClickHouse cluster deployment with Altinity operator or ClickHouse Operator
- Shard and replica topology design for multi-factory deployments
- Storage configuration: tiered storage,
disk policies, S3-backed cold storage
- Backup and restore strategies
- ClickHouse Keeper vs ZooKeeper for coordination
Observability & Monitoring
- Full observability stack deployment: Prometheus, Grafana, Alertmanager
- Kafka metrics: broker JMX metrics, consumer lag (Burrow or Cruise Control), under-replicated partitions
- Flink metrics: checkpoint duration, backpressure, restart frequency, throughput/latency per operator
- ClickHouse metrics: query latency, merge backlog, insert rate, replication lag, memory/disk pressure
- Kubernetes cluster metrics: node utilization, pod restarts, OOMKill events, PVC utilization
- SLO/SLI definition and alerting for platform uptime and data freshness
- Distributed tracing for pipeline end-to-end latency visibility
- Log aggregation (EFK stack, Loki, or equivalent) for platform component logs
Load Testing & Capacity Planning
- Load testing frameworks for streaming platforms (Gatling, k6, custom Kafka producer harnesses)
- Designing load test scenarios that simulate factory production volumes and burst conditions
- Benchmarking Kafka throughput (MB/s, msg/s) under varying partition counts and consumer configurations
- Flink job stress testing — measuring backpressure onset, checkpoint interval degradation under load
- ClickHouse ingestion benchmarking (insert throughput, concurrent query performance under load)
- Capacity modeling: translating factory data volumes and growth projections into infrastructure sizing
- Breaking point analysis — identifying bottlenecks before production traffic hits them
Isolation & Multi-Tenancy
- Namespace-level isolation for multiple factory deployments on shared Kubernetes clusters
- Kafka topic and consumer group isolation strategies across factory sites and teams
- Resource quotas and LimitRanges to prevent noisy-neighbor problems between factory workloads
- Network policy enforcement to isolate factory data plane traffic
- Separate Kafka clusters vs shared clusters with namespace isolation — trade-off analysis and implementation
- Per-factory ClickHouse cluster deployment vs shared cluster with database/user isolation
CI/CD & GitOps for Platform
- GitOps-driven platform deployment: ArgoCD or Flux for Kubernetes manifest management
- Helm chart versioning and promotion across environments (dev → staging → factory-prod)
- Operator upgrade management and CRD migration procedures
- Automated smoke tests post-deployment to validate platform health
- Change management for platform updates in production factory environments
Security & Compliance
- mTLS and certificate management (cert-manager on Kubernetes)
- Kafka ACL management at scale
- Network segmentation between factory OT and IT data flows
- Image scanning and supply chain security for platform container images
- Audit logging for platform access and configuration changes
AI-Assisted Development
- Using AI coding assistants (GitHub Copilot, Cursor, or equivalent) to accelerate Helm chart authoring, operator configuration, and runbook generation
- AI-assisted troubleshooting — using LLMs to diagnose Kubernetes events, Flink exceptions, and Kafka error logs faster
- Generating observability dashboards and alert rules with AI assistance
- Critical review of AI-generated infrastructure code before applying to production
📌 Platform Engineer (Bengaluru)
🏢 Dhruv Technology Solutions
📍 Bengaluru