As a Data Engineer, you will be responsible for designing, developing, and maintaining data solutions for data generation, collection, and processing in Big Data environments using predominantly PySpark/Python. The role involves building scalable data pipelines, ensuring data quality, and implementing ETL processes to support enterprise data platforms.
Key Responsibilities
- Design, develop, and maintain robust, scalable, high-performance data pipelines using PySpark.
- Create and manage ETL processes for data migration and deployment across systems.
- Migrate Ab Initio ETL applications into PySpark-based data pipelines.
- Migrate on-premise workloads to cloud platforms such as AWS, Databricks, and Snowflake.
- Collaborate with cross-functional teams to resolve data-related issues.
- Ensure data quality, reliability, and performance of data solutions.
- Stay updated with emerging data engineering technologies and best practices.
Required Skills
- Strong experience in PySpark and Python
- Hands-on experience with Hadoop Ecosystem
- Experience handling large-scale data processing and transformation
- AWS Cloud services
- Databricks and/or Snowflake
- Apache Airflow
- CI/CD pipelines and Git
- Solid debugging and problem-solving skills
- Excellent communication and collaboration skills
Good to Have
- AWS EKS
- Docker & Containers
If you are interested, please share your updated CV!