We are looking for an experienced Data Engineer with strong hands-on experience in Python, PySpark, SQL, Unix/Linux, Apache Airflow and Apache Iceberg. The candidate will be responsible for designing, developing and maintaining scalable data pipelines and data processing solutions.
Key Responsibilities
- Develop scalable data engineering solutions using Python and PySpark.
- Design and implement ETL/ELT data pipelines.
- Work extensively with Apache Iceberg for modern data lake/lakehouse implementations.
- Create and maintain Iceberg tables, schemas and partition strategies.
- Perform data ingestion, transformation and processing using PySpark.
- Write complex and optimized SQL queries.
- Develop and manage workflows using Apache Airflow.
- Create, schedule, monitor and troubleshoot Airflow DAGs.
- Work with Unix/Linux environments for development, execution and troubleshooting.
- Implement data quality, validation and reconciliation processes.
- Optimize PySpark jobs, SQL queries and data pipelines for performance.
- Troubleshoot production pipeline failures and resolve issues within SLA.
- Work with large-scale datasets and distributed data processing frameworks.
- Collaborate with data architects, analysts, developers and business teams.
- Participate in development, testing, deployment and production support.
- Follow coding, documentation, data security and data engineering best practices.
Mandatory Skills
- Python
- PySpark
- SQL
- Unix / Linux
- Apache Airflow
- Apache Iceberg
- Solid understanding of ETL/ELT
- Strong understanding of Data Lake / Data Lakehouse concepts
- Experience working with large-scale data processing
Apache Iceberg – Required Experience Candidates should have practical experience with:
- Apache Iceberg table management
- Table creation and schema evolution
- Partitioning and partition evolution
- Time travel
- Snapshot management
- Data ingestion and transformation
- Iceberg integration with Spark/PySpark
- Performance optimization
- Data lakehouse architecture
Good to Have
- Apache Spark
- Databricks
- AWS / Azure / GCP
- Hadoop / Hive
- Kafka
- Delta Lake / Hudi
- Snowflake
- AWS S3 / Azure Data Lake
- Git
- CI/CD
- Data warehouse concepts
- Cloud data engineering
Technical Skill Matrix SkillRequirementPython MandatoryPySpark MandatorySQL MandatoryUnix/Linux MandatoryApache Airflow MandatoryApache Iceberg MandatoryApache SparkGood to HaveData Lake / Lakehouse MandatoryCloudGood to HaveKafkaGood to HaveDatabricksGood to Have
📌 Python Developer Lead (Bengaluru)
🏢 Glauben Technologies
📍 Bengaluru
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.