Responsibility:
- Design, build, and maintain scalable data pipelines using AWS Glue and Databricks.
- Develop and optimize ETL/ELT workflows using PySpark and Python.
- Implement data ingestion frameworks for structured and unstructured data.
- Build and manage Data Lake/Lakehouse solutions on AWS.
- Create reusable data processing frameworks and data quality checks.
- Monitor, troubleshoot, and optimize data pipelines for performance and reliability.
- Collaborate with business stakeholders, architects, and development teams to gather requirements and deliver solutions.
- Ensure adherence to security, governance, and data management best practices.
- Participate in code reviews and mentor junior data engineers.
- Support production deployments and resolve critical issues.
Skills:
- Strong hands-on experience in AWS Glue
- Expertise in Databricks (Notebooks, Jobs, Delta Lake)
- Advanced programming skills in PySpark and Python
- Robust SQL and Data Warehousing concepts
- Experience with AWS services such as S3, EMR, Lambda, IAM, Athena, Redshift
- Experience in designing and developing ETL/ELT pipelines
- Data Lake and Lakehouse architecture implementation
- Git/GitHub for version control
- Performance tuning and optimization of Spark jobs