Our client is seeking a Data Engineer to build high-performance pipelines that power their Explainable AI engine. In this role, you will be responsible for ingesting, normalizing, and modeling complex billing metadata from AWS, Azure, GCP, and Snowflake, ensuring that autonomous agents operate on a foundation of high-fidelity data. A FinOps background is not required, as the domain is learnable, and the team will provide the necessary ramp-up.
Key Responsibilities
- Ingest raw structured and semi-structured data from customer cloud environments (AWS, Azure, GCP) into Snowflake staging layers.
- Configure ADF Linked Services, Integration Runtimes, and pipeline triggers for reliable data movement across multi-cloud boundaries.
- Optimize pipeline performance for high-volume datasets, ensuring low-latency delivery to downstream transformation layers.
- Pull data from REST APIs and third-party data sources as required by product needs.
- Transform and normalize API responses using Python and SQL to produce clean, structured tables in Snowflake.
- Ensure transformed data is structured and optimized for direct consumption by front-end applications and BI tools.
- Build reusable Python utilities for API pagination, error handling, and incremental data extraction.
- Collaborate with product and domain experts to gather and analyze business requirements for current product modules.
- Translate business requirements into clean, scalable data models.
- Design analytics-ready schemas (Star Schema, dimensional models) for ML and front-end consumption.
- Implement data quality frameworks to ensure accuracy, completeness, and consistency across all pipeline layers.
- Propose refresh strategies for each table, identifying appropriate SCD types.
- Implement automated testing and watermarking logic to prevent data gaps in Financial Reporting.
- Instrument observability metadata for monitoring data quality and performance.