Job Description (10 Key Responsibilities)
Design, develop, and maintain scalable ETL/ELT pipelines using AWS Glue and PySpark.
Ingest, transform, and load structured and semi-structured data from multiple sources into Snowflake.
Develop and optimize complex SQL queries, stored procedures, and views for analytics and reporting.
Build and manage data workflows using Glue Workflows, Triggers, and AWS Step Functions.
Implement data partitioning and performance tuning techniques in Snowflake and AWS settings.
Work with S3, Lambda, IAM, CloudWatch, and other AWS services to ensure secure and reliable pipelines.
Perform data validation, reconciliation, and quality checks across source and target systems.
Collaborate with data analysts, BI developers, and stakeholders to understand data requirements.
Automate deployment using CI/CD pipelines and Infrastructure-as-Code where applicable.
Monitor and troubleshoot data pipelines and ensure SLA adherence.