- Design, deploy, and manage Kafka clusters across multiple data centers to ensure high availability and scalability.
- Troubleshoot issues with Kafka brokers, topics, partitions, and consumer/producer applications.
- Collaborate with development teams to design and implement messaging solutions using Apache Kafka.
- Monitor Kafka cluster performance metrics such as throughput, latency, and resource utilization.
Job Requirements :
- 4-15 years of experience in managing large-scale distributed systems like Apache Kafka.
- Solid understanding of cloud-based infrastructure management using Terraform or similar tools.
- Experience with monitoring tools like Prometheus/Grafana for real-time visibility into system performance.