What makes this role different
Most data engineering ends at a table. Pattern's ad-tech output leaves the warehouse and spends a client's advertising budget within the hour. A silently wrong join or an unguarded backfill is a customer-facing incident, not a dashboard discrepancy - so correctness, idempotency and data-quality gating are the job, not paperwork after the job.
n
The system you'll work on
Destiny is Pattern's automated Ads optimizer. Once a day it discovers the keywords worth buying for every eligible product, assembles a wide feature store from performance, bid-history and search-results data, runs 5 machine-learning models, and picks the bid level that hits each product group's return on ad spend (ROAS) and budget target. A second pipeline then pushes those campaign, keyword and budget edits to the marketplace Ads API every 15 minutes.
It is a large, opinionated data system: a roughly 17,000-line orchestrated SQL codebase, a feature and label store several hundred columns wide, 5 model training and batch-scoring jobs,
and blocking data-quality gates in front of every outward write. You would be one of the engineers who owns it end to end.
Roles and Responsibilities
•
Develop, deploy, and support automated, scalable batch data pipelines from a variety of sources into the lakehouse.
•
Own and extend Airflow orchestration for a multi-DAG, cross-triggered daily pipeline and a 15-minute action pipeline - including branching, parallel task groups, cross-DAG triggers, backfill and full-refresh paths, and secure reruns.
•
Write and tune large analytical SQL: multi-hundred-column joins, window functions, incremental merges, and the warehouse-sizing and query-profile work needed to keep a daily run inside its window and its budget.
•
Extend the feature store - add current features and labels, wire them through the join layer, and preserve the leakage and data-completeness conventions that make the models trainable.
•
Orchestrate model training and batch infe
📌 Senior Data Engineer Pune
🏢 Pattern
📍 Pune