04 Sep
|
STYLI
|
Bengaluru
Location: Bangalore (Yamlur)
Experience: 6-10 Years
Role: Principal Data Engineer
About Styli Marketplace
Launched in 2019 by Landmark Group, Styli Marketplace is the first e-commerce venture of the group, quickly becoming a leading online destination for fashion and lifestyle across the GCC, including Saudi Arabia, the UAE, Kuwait, Bahrain, and beyond. Styli connects global sellers and creators with millions of fashion-forward customers, offering the latest trends, exceptional value, and convenient services like same-day to 48-hour delivery and flexible payment options. Our mission is to make style accessible, aspirational, and exciting for all, backed by a passionate team fostering a culture of creativity and innovation.
Role Overview
We are looking for a Principal Data Engineer to take technical ownership of Styli's Data Platform and to lead its evolution from a BigQuery-centric warehouse to a modern Lakehouse on Databricks. This is a hands-on architecture and leadership role: you will set the technical direction for the platform, drive the migration end-to-end (design, staffing, partner management, cutover), and mentor a team of senior and staff data engineers. You will work closely with Platform/DevOps, Security, ML, and business stakeholders across the GCC and India to ensure the platform scales reliably, cost-efficiently, and securely as Styli grows.
What You'll Do
Data Platform Architecture & Migration Leadership
- Own the end-to-end target-state architecture for Styli's data platform, leading the migration from BigQuery to a Databricks Lakehouse built on Delta Lake / Apache Iceberg, Unity Catalog, and GCS.
- Lead build-vs-buy and platform evaluations (e.g., Databricks vs. Snowflake), and translate the recommendation into an executable, phased migration roadmap with clear cutover and rollback criteria.
- Define and defend architecture decision records (ADRs) covering table format strategy, catalog ownership, and cross-engine interoperability between Databricks and existing warehouses via the Iceberg REST catalog.
- Own the staffing and delivery model for the migration — partnering with Databricks implementation/SI partners, structuring hybrid onsite/offshore teams, and tracking cost, timeline, and risk.
- Drive Unity Catalog adoption for centralized governance, fine-grained access control, and data lineage across workspaces and markets.
Data Pipeline Development
- Design, build, and maintain batch and real-time pipelines migrating from Airflow/BigQuery patterns to Databricks Workflows and Delta Live Tables, alongside Apache Airflow and dbt where appropriate.
- Handle ingestion from diverse sources: MySQL/PostgreSQL, Kafka event streams, S3/GCS data lakes, REST APIs, and SaaS platforms.
- Ensure pipelines are fault-tolerant, recoverable, and instrumented with data quality checks at every stage, with clear migration parity/reconciliation testing against legacy BigQuery pipelines.
Data Modelling & Lakehouse Engineering
- Design and own the medallion (bronze/silver/gold) data architecture on Delta Lake,
including dimensional and analytical models (star schema, OBT, wide tables) for BI and ML consumption.
- Own warehouse/lakehouse performance engineering — partitioning, Z-ordering/clustering, Photon acceleration, and cost-efficient query and compute (cluster policy) patterns on Databricks, alongside the legacy BigQuery estate during transition.
- Set and enforce standards for dbt models, tests, documentation, and lineage across a well-governed transformation layer.
Streaming & Real-Time Data
- Build and maintain real-time pipelines for live inventory updates, order event processing, and customer behavior streams, using Databricks Structured Streaming, Apache Flink, Spark Streaming, or ksqlDB.
- Design Kafka topic schemas, partition strategies, and consumer group management for high-throughput e-commerce event flows.
Data Lake & Cloud Infrastructure
- Build and manage a well-structured lakehouse on GCS, with Databricks running natively on GCP alongside existing cloud-native data services.
- Partner with the Platform/DevOps team on infrastructure-as-code (Terraform) for data infrastructure provisioning and containerized data workloads on Kubernetes.
Data Quality & Observability
- Set the data quality framework (Outstanding Expectations, dbt tests, or Monte Carlo) to catch anomalies, nulls, duplicates, and schema drift before they reach downstream consumers — with particular rigor during migration cutover windows.
- Own pipeline SLAs, freshness SLOs, and schema-change alerting as first-class platform metrics; own incident response for data pipeline failures with transparent runbooks and escalation paths.
ML & Analytics Enablement
- Build and maintain feature pipelines for ML models powering personalisation, recommendations, fraud detection, and demand forecasting, migrating toward Databricks Feature Store / MLflow where it improves training–serving consistency.
- Enable self-serve analytics by maintaining clean semantic layers and well-documented data marts for BI tools (Looker, Metabase, Superset).
Team Leadership & Mentorship
- Set technical standards and best practices for the data engineering team; review designs and code for scalability, cost, and reliability.
- Mentor senior and mid-level data engineers, including team members based in Bengaluru, and act as the technical escalation point for the platform.
- Partner directly with Platform Engineering, Security, and business stakeholders to align the migration roadmap with compliance and cost objectives.
What We're Looking For
Required
- 6-10 years of hands-on data engineering experience, including demonstrated ownership of platform-level architecture, ideally in a high-transaction e-commerce, fintech,
or consumer tech environment.
- Proven, hands-on experience leading a large-scale migration onto Databricks — including Delta Lake, Unity Catalog, Databricks Workflows/Delta Live Tables, and cluster/cost optimization.
- Strong proficiency in SQL — complex analytical queries, window functions, CTEs, query optimisation, and warehouse-specific dialects (BigQuery / Databricks SQL / Redshift).
- Deep experience architecting lakehouse platforms on GCP, including evaluating and integrating open table formats (Apache Iceberg or Delta Lake) and cross-engine interoperability patterns.
- Hands-on experience with Apache Airflow for orchestration and dbt for transformation layer management, at scale — handling millions of events per day with reliability and observability.
- Working knowledge of Apache Kafka or equivalent event streaming platforms for real-time data ingestion.
- Proficiency in Python for pipeline development, data transformation, and automation, including PySpark for distributed processing on Databricks.
- Track record of technical mentorship and cross-functional stakeholder management, including partner/vendor (SI) management for large migration programs.
- Understanding of data modelling principles: normalisation, dimensional modelling, medallion architecture, and analytical patterns.
Good to Have
- Experience running Databricks and BigQuery (or Snowflake) in parallel during a phased migration, including reconciliation and cutover strategy design.
- Familiarity with data cataloguing and governance tooling (DataHub, Amundsen, Collibra, or Alation) beyond Unity Catalog.
- Knowledge of data mesh or data product principles and federated ownership models.
- Experience with real-time analytics databases: ClickHouse, Druid, or Pinot for high-throughput OLAP workloads.
- Familiarity with infrastructure-as-code (Terraform) and deploying data workloads on Kubernetes.
- Relevant certifications: Databricks Certified Data Engineer Professional, Databricks Lakehouse Architect, Google Professional Data Engineer, or dbt Analytics Engineering.
The Data Problems You'll Solve at Styli
- Leading the BigQuery-to-Databricks lakehouse migration — architecting and executing the move for a live, high-traffic e-commerce platform with zero-downtime tolerance.
- Personalisation at scale — building event pipelines that capture every click, view, and purchase to power real-time recommendations for millions of customers.
- Inventory & demand forecasting — reliable pipelines feeding ML models that optimise stock levels across hundreds of SKUs and multiple markets.
- Campaign & marketing analytics — ingesting data from paid channels, CRM, and app attribution to give growth teams a single source of truth.
- Order & payments intelligence — real-time event streams for order lifecycle tracking, fraud signals, and settlement reconciliation.
- Cross-market reporting — unified data models serving multiple country markets (GCC and India) with different currencies, catalogues, and fulfilment providers.
📌 Principal Data Engineer (Bengaluru)
🏢 STYLI
📍 Bengaluru