- Ingest data from various internal and external sources into AWS Redshift and S3 buckets using methods such as DMS, Zero-ETL, Kinesis, Glue, Lambda, Lake Formation, Cross Account Replication, and SFTP.
- Build infrastructure as code using AWS CDK.
- Create and maintain GitLab CI/CD pipelines for code promotion through test and production environments.
- Design, build, and maintain AWS Glue ETL pipelines for structuring and curating data.
- Conduct code reviews on all deployments to ensure quality and best practices.
- Maintain AWS environments to optimize costs, reduce vulnerabilities, and ensure smooth operations.
- Coordinate or resolve production issues in a timely manner.
- Drive path-to-production processes,
including proper documentation and approvals.
- Collaborate with business teams on data governance and compliance initiatives.
- Identify and implement best practices for data ingestion, data design, and data quality improvement.
- Develop queries for profiling data, validating analyses, testing assumptions, and driving data quality assessment.
- Optimize query performance using indexing, materialized views, and other techniques.