30 Sep
|
Affinity Global
|
Sangli
30 Sep
Affinity Global
Sangli
About Affinity Affinity is pioneering new frontiers in AdTech: developing solutions that push past today’s limits and open up new opportunities. We are a global AdTech company helping publishers discover better ways to monetize and enabling advertisers to reach the right audiences through new touchpoints. Operating across 10+ markets in Asia, the US, and Europe with a team of over 500 experts, we are building privacy-first ad infrastructure that opens up opportunities beyond the walled gardens.
Team Lead, Data Engineer Work Location: Mumbai (Malad) Experience: 10+ Years About Role: We are looking for a Sr. Data Engineer to join the Core Architecture team and take ownership of the data platform end-to-end — from schema design and data ingestion to processing, storage and serving. The role will work across large-scale data systems involving Apache Spark, ClickHouse/MPP data stores, streaming pipelines, table formats and self-managed infrastructure.
The ideal candidate should have solid hands-on experience building and tuning production data platforms where performance, reliability and cost efficiency are critical. This is a hands-on technical leadership role requiring someone who can make architecture and engineering decisions, troubleshoot complex data-platform performance issues, operate infrastructure at scale and mentor engineers.
Roles & Responsibility: · Own the end-to-end data platform architecture covering schema design, ingestion, processing, storage and data serving. · Design and develop highly scalable batch and streaming data pipelines using Apache Spark. · Build, tune and optimise Spark jobs for performance, throughput, resource utilisation and reliability. · Design and optimise ClickHouse or equivalent MPP/columnar data platforms for high-volume analytical workloads. · Own table design, partitioning, sorting/indexing strategies, materialized views and query optimisation. · Troubleshoot and optimise database performance including query plans, partition pruning, high-cardinality workloads, compaction and defragmentation.
· Manage high-ingest data workloads and optimise storage/compute performance under fixed infrastructure capacity. · Work with modern table formats such as Apache Iceberg, Delta Lake or Apache Hudi, including schema evolution, compaction and small-file management. · Design reliable ingestion and streaming architectures for high-volume data processing. · Own production data-platform reliability, availability and performance, particularly in self-managed infrastructure environments. · Design systems with a robust focus on capacity planning, resource utilisation and cost optimisation, rather than relying on continuous infrastructure scaling. · Work across AWS/GCP/Azure environments and integrate cloud services with self-managed data infrastructure. · Troubleshoot complex production issues across Spark, databases, streaming pipelines, storage and infrastructure layers. · Establish engineering standards for data-platform design, performance, reliability and operational excellence. · Evaluate and implement technologies that improve scalability, performance and cost efficiency. · Collaborate with the Core Architecture, Platform, Infrastructure and Engineering teams on the evolution of the Data platform. · Mentor engineers and contribute to building and strengthening the data engineering sub-team. · Support hiring and technical evaluation of data engineering talent as the team expands.
Required Skills
- Experience: 10+ years in Data Engineering/Data Platform Engineering with end-to-end platform ownership.
- Core Tech Stack: Hands-on Apache Spark (batch/streaming) and ClickHouse (or equivalent MPP/columnar DBs).
- Data Storage & formats: Expertise in open table formats (Iceberg, Delta Lake, or Hudi) and database performance tuning (partitioning, compaction, query optimization).
- Infrastructure & Cloud: Operating self-managed data infrastructure alongside major cloud platforms (AWS, GCP, or Azure) under fixed-capacity/cost constraints.
- Languages & Core Skills: Advanced SQL, data modeling, and Scala, Java, or Python.
- Nice-to-Haves / Pluses: Kafka, Flink, Kafka Streams, and CDC pipelines (e.g., Debezium)
📌 Team Lead, Data Engineer (Sangli)
🏢 Affinity Global
📍 Sangli