POSITION SUMMARY : We are seeking an experienced AI Data Engineer to build andoptimizedata pipelines that power AI/ML solutions across ETS. This role contributes to data collection efforts through the design and deployment of data instruments, oversees aggregation and cleaning processes, and collaborates with data scientists and ML engineers to interpret and implement findings. The position is critical to preparing high-quality, scalable, and compliant data for advanced analytics, automation, and machine learning use cases while supporting data visualization and trend analysis efforts.
PRIMARY RESPONSIBILITIES
Design and implement robust, scalable data pipelines to support AI/ML model development and deployment
Clean, transform, and curate structured and unstructured data from diverse sources to ensure model-ready quality
Collaborate with data scientists, ML engineers, and business teams to ensure data readiness, usability,
and alignment with AI objectives
Develop andmaintainmetadata management, data lineage, and data quality frameworks to support AI governance and compliance
Enable advanced feature engineering capabilities and implement real-time data streaming solutions for AI applications
Design and deploy data collection instruments and oversee data aggregation processes
Ensure data compliance, privacy, and ethical use standards across all AI workflows
Support enterprise-wide data initiatives including business glossary development, taxonomy creation, and the DARE program''s automation goals
Must Have -
Python advanced SQL
Spark / distributed processing
Cloud platforms (AWS)
Airflow / orchestration tools
Data pipeline automation
Data governance quality
Nice to have -
Kafka/Kinesis streaming
Feature stores
LLMready data pipelines
MLOps exposure