25 Sep
|
Affinity Global
|
Mumbai
25 Sep
Affinity Global
Mumbai
About Affinity
Affinity is pioneering new frontiers in AdTech: developing solutions that push past today’s limits and open up new opportunities. We are a global AdTech company helping publishers discover better ways to monetize and enabling advertisers to reach the right audiences through new touchpoints. Operating across 10+ markets in Asia, the US, and Europe with a team of over 500 experts, we are building privacy-first ad infrastructure that opens up opportunities beyond the walled gardens.
Team Lead, Data Engineer - DMP
Work Location: Mumbai (Malad)
Experience: 10+ Years
About Role:
We are looking for a Sr. Data Engineer, DMP to join the Core Architecture team and take ownership of the data platform end-to-end — from schema design and data ingestion to processing, storage and serving.
The role will work across large-scale data systems involving Apache Spark, ClickHouse/MPP data stores, streaming pipelines, table formats and self-managed infrastructure. The ideal candidate should have strong hands-on experience building and tuning production data platforms where performance, reliability and cost efficiency are critical.
This is a hands-on technical leadership role requiring someone who can make architecture and engineering decisions, troubleshoot complex data-platform performance issues, operate infrastructure at scale and mentor engineers.
Roles & Responsibility:
· Own the end-to-end data platform architecture covering schema design, ingestion, processing, storage and data serving.
· Design and develop highly scalable batch and streaming data pipelines using Apache Spark.
· Build, tune and optimise Spark jobs for performance, throughput, resource utilisation and reliability.
· Design and optimise ClickHouse or equivalent MPP/columnar data platforms for high-volume analytical workloads.
· Own table design, partitioning, sorting/indexing strategies, materialized views and query optimisation.
· Troubleshoot and optimise database performance including query plans, partition pruning, high-cardinality workloads, compaction and defragmentation.
· Manage high-ingest data workloads and optimise storage/compute performance under fixed infrastructure capacity.
· Work with up-to-date table formats such as Apache Iceberg, Delta Lake or Apache Hudi, including schema evolution, compaction and small-file management.
· Design reliable ingestion and streaming architectures for high-volume data processing.
· Own production data-platform reliability, availability and performance, particularly in self-managed infrastructure environments.
· Design systems with a strong focus on capacity planning, resource utilisation and cost optimisation, rather than relying on continuous infrastructure scaling.
· Work across AWS/GCP/Azure environments and integrate cloud services with self-managed data infrastructure.
· Troubleshoot complex production issues across Spark, databases, streaming pipelines, storage and infrastructure layers.
· Establish engineering standards for data-platform design, performance, reliability and operational excellence.
· Evaluate and implement technologies that improve scalability, performance and cost efficiency.
· Collaborate with the Core Architecture, Platform, Infrastructure and Engineering teams on the evolution of the DMP platform.
· Mentor engineers and contribute to building and strengthening the data engineering sub-team.
· Support hiring and technical evaluation of data engineering talent as the team expands.
Required Skills:
- Experience: 10+ years in Data Engineering/Data Platform Engineering with end-to-end platform ownership.
- Core Tech Stack: Hands-on Apache Spark (batch/streaming) and ClickHouse (or equivalent MPP/columnar DBs).
- Data Storage & formats: Expertise in open table formats (Iceberg, Delta Lake, or Hudi) and database performance tuning (partitioning, compaction, query optimization).
- Infrastructure & Cloud: Operating self-managed data infrastructure alongside major cloud platforms (AWS, GCP, or Azure) under fixed-capacity/cost constraints.
- Languages & Core Skills: Advanced SQL, data modeling, and Scala, Java, or Python.
- Nice-to-Haves / Pluses: Kafka, Flink, Kafka Streams, and CDC pipelines (e.g., Debezium)
📌 Team Lead, Data Engineer - DMP (Mumbai)
🏢 Affinity Global
📍 Mumbai