Principal Data Engineer (Bengaluru)

Principal Data Engineer (Bengaluru)

04 Sep
|
STYLI
|
Bengaluru

04 Sep

STYLI

Bengaluru

Location: Bangalore (Yamlur)

Experience: 6-10 Years

Role: Principal Data Engineer

About Styli Marketplace

Launched in 2019 by Landmark Group, Styli Marketplace is the first e-commerce venture of the group, quickly becoming a leading online destination for fashion and lifestyle across the GCC, including Saudi Arabia, the UAE, Kuwait, Bahrain, and beyond. Styli connects global sellers and creators with millions of fashion-forward customers, offering the latest trends, exceptional value, and convenient services like same-day to 48-hour delivery and flexible payment options. Our mission is to make style accessible, aspirational, and exciting for all, backed by a passionate team fostering a culture of creativity and innovation.

Role Overview

We are looking for a Principal Data Engineer to take technical ownership of Styli's Data Platform and to lead its evolution from a BigQuery-centric warehouse to a modern Lakehouse on Databricks. This is a hands-on architecture and leadership role: you will set the technical direction for the platform, drive the migration end-to-end (design, staffing, partner management, cutover), and mentor a team of senior and staff data engineers. You will work closely with Platform/DevOps, Security, ML, and business stakeholders across the GCC and India to ensure the platform scales reliably, cost-efficiently, and securely as Styli grows.

What You'll Do

Data Platform Architecture & Migration Leadership

- Own the end-to-end target-state architecture for Styli's data platform, leading the migration from BigQuery to a Databricks Lakehouse built on Delta Lake / Apache Iceberg, Unity Catalog, and GCS.
- Lead build-vs-buy and platform evaluations (e.g., Databricks vs. Snowflake), and translate the recommendation into an executable, phased migration roadmap with clear cutover and rollback criteria.
- Define and defend architecture decision records (ADRs) covering table format strategy, catalog ownership, and cross-engine interoperability between Databricks and existing warehouses via the Iceberg REST catalog.
- Own the staffing and delivery model for the migration — partnering with Databricks implementation/SI partners, structuring hybrid onsite/offshore teams, and tracking cost, timeline, and risk.
- Drive Unity Catalog adoption for centralized governance, fine-grained access control, and data lineage across workspaces and markets.

Data Pipeline Development

- Design, build, and maintain batch and real-time pipelines migrating from Airflow/BigQuery patterns to Databricks Workflows and Delta Live Tables, alongside Apache Airflow and dbt where appropriate.
- Handle ingestion from diverse sources: MySQL/PostgreSQL, Kafka event streams, S3/GCS data lakes, REST APIs, and SaaS platforms.
- Ensure pipelines are fault-tolerant, recoverable, and instrumented with data quality checks at every stage, with clear migration parity/reconciliation testing against legacy BigQuery pipelines.

Data Modelling & Lakehouse Engineering

- Design and own the medallion (bronze/silver/gold) data architecture on Delta Lake,



including dimensional and analytical models (star schema, OBT, wide tables) for BI and ML consumption.
- Own warehouse/lakehouse performance engineering — partitioning, Z-ordering/clustering, Photon acceleration, and cost-efficient query and compute (cluster policy) patterns on Databricks, alongside the legacy BigQuery estate during transition.
- Set and enforce standards for dbt models, tests, documentation, and lineage across a well-governed transformation layer.

Streaming & Real-Time Data

- Build and maintain real-time pipelines for live inventory updates, order event processing, and customer behavior streams, using Databricks Structured Streaming, Apache Flink, Spark Streaming, or ksqlDB.
- Design Kafka topic schemas, partition strategies, and consumer group management for high-throughput e-commerce event flows.

Data Lake & Cloud Infrastructure

- Build and manage a well-structured lakehouse on GCS, with Databricks running natively on GCP alongside existing cloud-native data services.
- Partner with the Platform/DevOps team on infrastructure-as-code (Terraform) for data infrastructure provisioning and containerized data workloads on Kubernetes.

Data Quality & Observability

- Set the data quality framework (Outstanding Expectations, dbt tests, or Monte Carlo) to catch anomalies, nulls, duplicates, and schema drift before they reach downstream consumers — with particular rigor during migration cutover windows.
- Own pipeline SLAs, freshness SLOs, and schema-change alerting as first-class platform metrics; own incident response for data pipeline failures with transparent runbooks and escalation paths.

ML & Analytics Enablement

- Build and maintain feature pipelines for ML models powering personalisation, recommendations, fraud detection, and demand forecasting, migrating toward Databricks Feature Store / MLflow where it improves training–serving consistency.
- Enable self-serve analytics by maintaining clean semantic layers and well-documented data marts for BI tools (Looker, Metabase, Superset).

Team Leadership & Mentorship

- Set technical standards and best practices for the data engineering team; review designs and code for scalability, cost, and reliability.
- Mentor senior and mid-level data engineers, including team members based in Bengaluru, and act as the technical escalation point for the platform.
- Partner directly with Platform Engineering, Security, and business stakeholders to align the migration roadmap with compliance and cost objectives.

What We're Looking For

Required

- 6-10 years of hands-on data engineering experience, including demonstrated ownership of platform-level architecture, ideally in a high-transaction e-commerce, fintech,



or consumer tech environment.
- Proven, hands-on experience leading a large-scale migration onto Databricks — including Delta Lake, Unity Catalog, Databricks Workflows/Delta Live Tables, and cluster/cost optimization.
- Strong proficiency in SQL — complex analytical queries, window functions, CTEs, query optimisation, and warehouse-specific dialects (BigQuery / Databricks SQL / Redshift).
- Deep experience architecting lakehouse platforms on GCP, including evaluating and integrating open table formats (Apache Iceberg or Delta Lake) and cross-engine interoperability patterns.
- Hands-on experience with Apache Airflow for orchestration and dbt for transformation layer management, at scale — handling millions of events per day with reliability and observability.
- Working knowledge of Apache Kafka or equivalent event streaming platforms for real-time data ingestion.
- Proficiency in Python for pipeline development, data transformation, and automation, including PySpark for distributed processing on Databricks.
- Track record of technical mentorship and cross-functional stakeholder management, including partner/vendor (SI) management for large migration programs.
- Understanding of data modelling principles: normalisation, dimensional modelling, medallion architecture, and analytical patterns.

Good to Have

- Experience running Databricks and BigQuery (or Snowflake) in parallel during a phased migration, including reconciliation and cutover strategy design.
- Familiarity with data cataloguing and governance tooling (DataHub, Amundsen, Collibra, or Alation) beyond Unity Catalog.
- Knowledge of data mesh or data product principles and federated ownership models.
- Experience with real-time analytics databases: ClickHouse, Druid, or Pinot for high-throughput OLAP workloads.
- Familiarity with infrastructure-as-code (Terraform) and deploying data workloads on Kubernetes.
- Relevant certifications: Databricks Certified Data Engineer Professional, Databricks Lakehouse Architect, Google Professional Data Engineer, or dbt Analytics Engineering.

The Data Problems You'll Solve at Styli

- Leading the BigQuery-to-Databricks lakehouse migration — architecting and executing the move for a live, high-traffic e-commerce platform with zero-downtime tolerance.
- Personalisation at scale — building event pipelines that capture every click, view, and purchase to power real-time recommendations for millions of customers.
- Inventory & demand forecasting — reliable pipelines feeding ML models that optimise stock levels across hundreds of SKUs and multiple markets.
- Campaign & marketing analytics — ingesting data from paid channels, CRM, and app attribution to give growth teams a single source of truth.
- Order & payments intelligence — real-time event streams for order lifecycle tracking, fraud signals, and settlement reconciliation.
- Cross-market reporting — unified data models serving multiple country markets (GCC and India) with different currencies, catalogues, and fulfilment providers.

📌 Principal Data Engineer (Bengaluru)
🏢 STYLI
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: principal data engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: principal data engineer (bengaluru) / bengaluru