28 Sep
|
Apexon
|
Bengaluru
Data Migration Factory Data Engineer
Experience -8 to 13 Years
About the Role
We are looking for a highly skilled Data Engineer to join a large-scale enterprise data modernization program focused on migrating legacy Hadoop Data Lake platforms to a contemporary Snowflake Lakehouse architecture. The role involves designing, developing, and executing data migration pipelines, reconciling migrated data, and ensuring seamless integration with target Lakehouse settings.
The ideal candidate should possess robust expertise in Python, SQL, Spark, Hadoop ecosystem technologies, and Snowflake, along with hands-on experience in building scalable ETL/ELT pipelines and data migration solutions.
Key Responsibilities
Analyze existing Hadoop-based data pipelines and migration requirements.
Design and develop migration frameworks for moving data from Hadoop Data Lake to Snowflake Lakehouse.
Rewrite and optimize existing Spark/Hadoop jobs for target architecture.
Develop one-time migration and incremental data load processes.
Build scalable ETL/ELT pipelines using Spark and Python.
Perform data validation, reconciliation, and quality checks throughout migration phases.
Work with structured, semi-structured,
and large-scale enterprise data sets.
Optimize data processing performance and resource utilization.
Collaborate with architects, business analysts, and downstream application teams.
Troubleshoot and resolve data quality, performance, and pipeline issues.
Create technical documentation, migration plans, and operational runbooks.
Support testing, deployment, and production migration activities.
Mandatory Skills
Core Technologies
Python
SQL
Apache Spark / PySpark
Hadoop Ecosystem (HDFS, Hive, Sqoop)
Data Engineering & ETL Development
Snowflake
Data Platforms
Hadoop Data Lake
Snowflake Lakehouse
Data Warehousing Concepts
Data Processing
Spark SQL
Data Transformation
Data Loading Frameworks
Data Validation & Reconciliation
Performance Optimization
Orchestration & Automation
Apache Airflow
Workflow Scheduling
Pipeline Automation
Preferred Skills
Apache Iceberg
Parquet File Formats
Databricks
AWS (S3, Glue, EMR, Lambda)
Azure Data Engineering Services
Kafka
Data Lineage Tools
Metadata Management Platforms
CI/CD for Data Pipelines
📌 Senior Data Engineer Bengaluru
🏢 Apexon
📍 Bengaluru