20 Sep
|
Apexon
|
Bengaluru
Data Migration Factory Data Engineer
Experience
3 to 8 Years (Minimum 3+ years of hands-on Data Engineering experience)
About the Role
We are looking for a highly skilled Data Engineer to join a large-scale enterprise data modernization program focused on migrating legacy Hadoop Data Lake platforms to a modern Snowflake Lakehouse architecture. The role involves designing, developing, and executing data migration pipelines, reconciling migrated data, and ensuring seamless integration with target Lakehouse environments.
The ideal candidate should possess solid expertise in Python, SQL, Spark, Hadoop ecosystem technologies, and Snowflake, along with hands-on experience in building scalable ETL/ELT pipelines and data migration solutions.
Key Responsibilities
- Analyze existing Hadoop-based data pipelines and migration requirements.
- Design and develop migration frameworks for moving data from Hadoop Data Lake to Snowflake Lakehouse.
- Rewrite and optimize existing Spark/Hadoop jobs for target architecture.
- Develop one-time migration and incremental data load processes.
- Build scalable ETL/ELT pipelines using Spark and Python.
- Perform data validation, reconciliation, and quality checks throughout migration phases.
- Work with structured,
semi-structured, and large-scale enterprise data sets.
- Optimize data processing performance and resource utilization.
- Collaborate with architects, business analysts, and downstream application teams.
- Troubleshoot and resolve data quality, performance, and pipeline issues.
- Create technical documentation, migration plans, and operational runbooks.
- Support testing, deployment, and production migration activities.
Mandatory Skills
Core Technologies
- Python
- SQL
- Apache Spark / PySpark
- Hadoop Ecosystem (HDFS, Hive, Sqoop)
- Data Engineering & ETL Development
- Snowflake
Data Platforms
- Hadoop Data Lake
- Snowflake Lakehouse
- Data Warehousing Concepts
Data Processing
- Spark SQL
- Data Transformation
- Data Loading Frameworks
- Data Validation & Reconciliation
- Performance Optimization
Orchestration & Automation
- Apache Airflow
- Workflow Scheduling
- Pipeline Automation
Preferred Skills
- Apache Iceberg
- Parquet File Formats
- Databricks
- AWS (S3, Glue, EMR, Lambda)
- Azure Data Engineering Services
- Kafka
- Data Lineage Tools
- Metadata Management Platforms
- CI/CD for Data Pipelines
📌 Data Engineer - Snowflake (Bengaluru)
🏢 Apexon
📍 Bengaluru