Responsibility:
Design, build, and maintain scalable data pipelines using AWS Glue and Databricks.
Develop and optimize ETL/ELT workflows using PySpark and Python.
Implement data ingestion frameworks for structured and unstructured data.
Build and manage Data Lake/Lakehouse solutions on AWS.
Create reusable data processing frameworks and data quality checks.
Monitor, troubleshoot, and optimize data pipelines for performance and reliability.
Collaborate with business stakeholders, architects, and development teams to gather requirements and deliver solutions.
Ensure adherence to security, governance, and data management best practices.
Participate in code reviews and mentor junior data engineers.
Support production deployments and resolve critical issues.
Skills:
Robust hands-on experience in AWS Glue
Expertise in Databricks (Notebooks, Jobs, Delta Lake)
Advanced programming skills in PySpark and Python
Robust SQL and Data Warehousing concepts
Experience with AWS services such as S3, EMR, Lambda, IAM, Athena, Redshift
Experience in designing and developing ETL/ELT pipelines
Data Lake and Lakehouse architecture implementation
Git/GitHub for version control
Performance tuning and optimization of Spark jobs