Lead Data Engineer (Mumbai)

Lead Data Engineer (Mumbai)

30 Sep
|
Sourcebae
|
Mumbai

30 Sep

Sourcebae

Mumbai

Lead Data EngineerExperience

6–8 YearsLocation:

Bengaluru / HybridEmployment Type

Full-TimeRole OverviewWe are looking for an experienced

Lead Data Engineer with strong hands-on expertise in real-time data streaming, event processing, CDC, data integration, and up-to-date data engineering .The candidate will be responsible for designing, developing, and maintaining high-performance, production-grade data pipelines using

Apache Flink, Apache Kafka, Debezium, CDC, ClickHouse, and Apache Airflow .This is a hands-on engineering role requiring practical experience in building high-volume, low-latency streaming applications, developing scalable data pipelines, troubleshooting distributed systems, and optimizing data processing workloads.The ideal candidate should be comfortable working across the complete data pipeline—from source systems and CDC ingestion through Kafka and Flink processing to analytical storage, APIs, dashboards, and downstream integrations .Key Responsibilities1. Real-Time Data EngineeringDesign, develop, and maintain real-time data pipelines using

Apache Flink and Apache Kafka .Develop production-grade streaming applications for high-volume and low-latency workloads.Implement data transformation, filtering, enrichment, aggregation, and event processing.Build reliable event-processing pipelines with appropriate error handling and recovery mechanisms.Consume and publish events across Kafka topics.Implement partitioning, consumer groups, offsets, and appropriate delivery mechanisms.Troubleshoot streaming pipeline failures, latency, throughput, and performance issues.2.

Apache

FlinkDevelop and maintain production-grade

Apache Flink jobs .Implement stream transformations, filtering, mapping, aggregations, joins, and windows.Work with event-time processing, watermarks, and state management .Implement Flink checkpointing, savepoints, and recovery mechanisms.Optimize Flink jobs for performance, scalability, and resource utilization.Monitor latency, throughput, failures, backpressure, and resource consumption.Troubleshoot state, checkpointing, backpressure, and processing issues.3.

Apache

KafkaDevelop Kafka-based ingestion and streaming pipelines.Create and manage Kafka topics and event streams.Work with partitions, offsets, consumer groups, replication, and retention .Develop reliable Kafka producer and consumer applications.Handle message ordering, retries, duplicate events,



and replay scenarios.Monitor Kafka performance and troubleshoot consumer lag and throughput issues.Work with Kafka schemas and serialization formats.4. CDC &

- DebeziumBuild CDC-based ingestion pipelines using

Debezium .Configure and maintain Debezium connectors.Capture source-system inserts, updates, and deletes.Publish CDC events into Kafka.Handle initial snapshots and incremental CDC processing.Manage schema evolution and source-system changes.Implement data reconciliation and consistency checks.Troubleshoot CDC failures and source-to-target data issues.5. Data OrchestrationDevelop and maintain data workflows using

Apache Airflow or equivalent orchestration frameworks .Build reusable DAGs for:Data ingestionCDC workflowsData validationFlink job executionData transformationClickHouse loadingDownstream integrationsImplement workflow dependencies, scheduling, retries, backfills, SLAs, and alerting.Integrate Airflow with Kafka, Flink, Debezium, ClickHouse, APIs, and cloud services.Monitor workflow execution and troubleshoot failures.Develop reusable operators, sensors, and workflow components where required.Use event-driven triggers for real-time workflows where appropriate.6. ClickHouse &

- Analytical DataIntegrate streaming data pipelines with

ClickHouse .Design efficient analytical data models.Develop and optimize SQL queries.Implement appropriate partitioning, sorting, indexing, and retention strategies.Optimize data ingestion and query performance.Support analytical use cases, dashboards, and reporting requirements.7. Data Quality &

- ReliabilityImplement data validation and quality checks throughout the data pipeline.Build reconciliation mechanisms between source and target systems.Monitor data freshness, completeness, accuracy, and consistency.Implement error handling, retry, replay, and recovery mechanisms.Establish logging and observability for critical pipelines.Support incident investigation and root-cause analysis.8. Integration &
- APIsIntegrate streaming and analytical data with

APIs, dashboards, endpoints, and downstream applications .Develop data interfaces and integration components.Work with application teams to define data contracts and integration requirements.Support future integrations and additional data consumers.9.

Engineering





PracticesFollow modern software engineering practices including:Git and version controlCode reviewsUnit and integration testingCI/CDLogging and monitoringDocumentationDevelop reusable, scalable, and maintainable data engineering components.Participate in technical design discussions and architecture reviews.Mentor Data Engineers and contribute to engineering standards.Required Skills &

- Experience6–8 years of experience in Data Engineering.Strong hands-on experience with Apache Flink – Mandatory/Core Requirement.Strong hands-on experience with

Apache Kafka .Hands-on experience with

Debezium and Change Data Capture (CDC) .Strong programming experience in

Java or Scala .Good experience with

Python is an advantage.Strong

SQL skills.Experience with analytical databases;

ClickHouse is highly preferred .Hands-on experience with

Apache Airflow or another data orchestration framework .Strong understanding of distributed systems and real-time data processing.Experience developing and supporting production-grade streaming pipelines.Strong understanding of Kafka concepts including:TopicsPartitionsOffsetsConsumer GroupsReplicationRetentionExperience with data transformation, enrichment, filtering, aggregation, and event processing.Strong troubleshooting and problem-solving skills for performance and reliability issues.Familiarity with cloud platforms and containerized environments.Preferred SkillsAdvanced Apache Flink experience, including:State ManagementCheckpointsSavepointsWatermarksEvent TimeWindowsBackpressureState BackendsExperience with

Apache Airflow, Dagster, Prefect, or Apache NiFi .Experience with

Kafka Schema Registry .Experience with

Avro, Protobuf, or JSON .Experience with

Kubernetes .Experience with

AWS, Azure, or GCP .Experience with

CI/CD pipelines .Experience with

Docker and containerized applications .Experience with

Terraform or other Infrastructure as Code tools .Experience with data observability and monitoring tools.Experience building high-volume, low-latency real-time data platforms.Experience with

REST APIs and system integrations .Key Technical StackApache Flink | Apache Kafka | Debezium | CDC | ClickHouse | Apache Airflow | Java/Scala | Python | SQL | Kubernetes | Docker | Cloud | CI/CD | REST APIs

Apply NowInterested candidates can share their updated CV at or WhatsApp it to

(phone hidden) .Stay updated with our latest job opportunities and company news by following us on LinkedIn:Sourcebae on LinkedIn

📌 Lead Data Engineer (Mumbai)
🏢 Sourcebae
📍 Mumbai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: lead data engineer (mumbai) / mumbai

Subscribe to this job alert:

Get the latest job offers by email for: lead data engineer (mumbai) / mumbai