Hiring a Data Engineer with solid hands-on experience in Spark/PySpark, Kafka, and SQL. The role focuses on building scalable ETL pipelines and working on large-scale data processing systems.
Key Responsibilities
Build and maintain ETL/ELT data pipelines
Develop data processing solutions using Spark/PySpark
Work with Kafka for data ingestion and processing
Ensure data quality and optimize performance
Collaborate with analytics and data science teams
Use orchestration tools such as Airflow, AWS Glue, or ADF
Required Skills (Key Requirements)
Mandatory:
Robust experience in Apache Spark / PySpark
Hands-on experience with Kafka
Advanced SQL skills
Good to Have:
Experience with ETL pipeline development
Knowledge of Apache Flink
Exposure to Airflow / ADF / AWS Glue
Experience with cloud platforms (AWS / Azure / GCP)
Understanding of data lakes and file formats (Parquet, Delta, etc.)
Important Note
This is not a Databricks-focused role
Candidates with solid Kafka + Spark + SQL experience will be prioritized