We need a Data Engineer for one of our US based clients. The project includes building a sensing and replenishment recommendation platform on AWS for a large haircare brand. The right candidate will work closely with the client on an application within AWS infrastructure. Core responsibility will include to build the ingestion pipeline and the curated layer that everything else depends on, then implements the deterministic cover and reorder logic.
Responsibilities
- Onboard seven Amazon data entities: sales, inventory, orders, product catalog, demand plan, promotional calendar, SKU-to-ASIN crosswalk
- Build the Delta Sharing ingestion framework — Glue 5.0 reading shared tables directly from the client's Unity Catalog and materializing them into the S3 Iceberg lakehouse. One reusable pattern, with entities onboarded as configuration rather than bespoke code
- Implement schema validation and reconciliation checks (the client provides data validated and clean, so this is verification rather than cleansing)
- Build the curated and semantic layer consumed by both the forecasting models and QuickSight
- Implement the deterministic cover projection and reorder point logic — on-hand plus inbound against run rate against a fixed lead time
- Implement min-threshold trigger logic and edge case handling
- Orchestrate the daily batch with Step Functions and EventBridge, with failure alerting
Essential experience
- 5+ years data engineering, with 3+ on AWS
- Solid PySpark and Python; production experience with AWS Glue (5.0 preferred)
- Apache Iceberg or Delta Lake table formats, partitioning and schema evolution
- SQL to a high standard — the cover and reorder logic is substantially SQL and business rules
- Step Functions, Lambda, EventBridge orchestration
- Experience building configuration-driven pipelines rather than one job per source
Desirable
- Databricks, Unity Catalog and Delta Sharing on the consuming side
- Experience with modular, tested, version-controlled transformation frameworks (dbt or equivalent)
- Prior work on inventory or supply chain data models