3+ years of experience in data engineering with hands-on Python, Spark, and SQL.
Strong proficiency in Python (data processing, scripting, OOP).
Hands-on experience with Apache Spark (PySpark) for large-scale data processing.
Advanced knowledge of SQL - joins, window functions, performance tuning.
Experience processing structured and semi-structured data from multiple sources.
Robust understanding of ETL/ELT concepts and best practices.
Good to have: experience with cloud platforms (AWS, Azure, GCP), data lakes (Hive, Delta Lake, Iceberg), and orchestration tools (Airflow).
Key Responsibilities*
Design, develop, and maintain scalable data pipelines using Python and Apache Spark.
Write optimized SQL queries for data extraction, transformation, and analysis.
Process structured and semi-structured data from multiple sources.
Optimize Spark jobs for performance and cost efficiency.
Implement data validation, error handling, and monitoring.
Collaborate with stakeholders to understand and fulfill data requirements.
Ensure data security, reliability, and adherence to best practices.