Roles and Responsibilities
Design, develop, test, deploy, and maintain large-scale data pipelines using AWS services such as S3, Glue, Lambda, Step Functions.
Collaborate with cross-functional teams to gather requirements and design solutions that meet business needs.
Develop complex SQL queries to extract insights from large datasets stored in Amazon Redshift and Athena.
Troubleshoot issues related to EMR clusters, Spark jobs, and data pipeline failures.
Ensure high availability and scalability of the system by implementing monitoring tools like CloudWatch.