Data Pipeline Development: Design, develop, and manage robust ETL/ELT pipelines to process structured and unstructured data efficiently. Ensure seamless data ingestion, transformation, and delivery across various systems. AWS Glue Expertise: Develop, schedule, and monitor Glue jobs for data extraction, transformation, and loading (ETL). Write and optimize PySpark scripts within Glue for large-scale data processing. Configure Glue crawlers for metadata cataloging and schema discovery. Troubleshoot and optimize Glue workflows for efficiency and cost-effectiveness.
Cloud Data Infrastructure: Build and maintain cloud-based data infrastructure using AWS services such as S3, Lambda, and Step Functions.
Automate data pipeline workflows using orchestration tools like Step Functions and AWS Glue workflows.
Database Management and SQL Expertise: Write, optimize, and maintain complex SQL queries for data extraction, reporting, and analytics. Administer relational and non-relational databases, ensuring data integrity and availability.
Preferred candidate profile
Work closely with data analysts, data scientists, and business stakeholders to understand and address data needs. Snowflake Expertise: Build data model in snowflake using complex SQL query Expert in GIT Hub for data pipeline deployment