As a Data Engineer, you will be responsible for designing, developing, and maintaining data solutions for data generation, collection, and processing in Big Data environments using predominantly PySpark/Python. The role involves building scalable data pipelines, ensuring data quality, and implementing ETL processes to support enterprise data platforms.
Key Responsibilities
Design, develop, and maintain robust, scalable, high-performance data pipelines using PySpark.
Create and manage ETL processes for data migration and deployment across systems.
Migrate Ab Initio ETL applications into PySpark-based data pipelines.
Migrate on-premise workloads to cloud platforms such as AWS, Databricks, and Snowflake.
Collaborate with cross-functional teams to resolve data-related issues.
Ensure data quality, reliability, and performance of data solutions.
Stay updated with emerging data engineering technologies and best practices.
Required Skills
Robust experience in PySpark and Python
Hands-on experience with Hadoop Ecosystem
Experience handling large-scale data processing and transformation
AWS Cloud services
Databricks and/or Snowflake
Apache Airflow
CI/CD pipelines and Git
Robust debugging and problem-solving skills
Excellent communication and collaboration skills
Valuable to Have
AWS EKS
Docker & Containers
If you are interested, please share your updated CV!