22 Sep
|
Bounteous
|
India
About This Role
6 to 9 years of experience
Skills: Kafka, ANSI SQL, FTP, Apache Spark , Snowflake, SQL, CI/CD pipelines, Python/Java
Role Overview
The Engineer will be part of the datastore-migration Factory team, responsible for the end-to-end datastore migration from an on-prem DataLake to an AWS-hosted LakeHouse. This is a high-visibility and crucial project.
Basic Qualifications
- Bachelor's or Master's in Computer Science, Applied Mathematics, Engineering, or a related quantitative field.
Technical Skills
- A total of 8+ years of experience in the field of Data Engineering.
- Minimum of 3–5 years of professional "hands-on-keyboard" coding experience in a collaborative, team-based environment.
- In-depth proficiency in SQL, Python/Java, and working experience using RESTful APIs with Swagger and automation; understand request/response and investigate using developer tools.
- Valuable knowledge of the on-prem Data Lake built on the Hadoop ecosystem with HDFS, and understanding of how the ecosystem operates.
- Experience with AWS S3 data staging and synchronous/asynchronous data sync techniques into a Lakehouse.
- Excellent experience handling CI/CD pipelines, stages, child pipeline management, investigating pipeline logs,
and root cause analysis.
- Ability to execute end-to-end migration of Legacy Data Lake data stores to a Lakehouse in the cloud using available tooling.
- Ability to investigate issues with tooling during data migration, perform root cause analysis, and report issues promptly to different collaborating tooling teams, tracking them to closure.
- Experience translating and optimizing legacy SQL and Spark-based consumption patterns (raw and modeled) for compatibility with Snowflake and Iceberg.
- Strong understanding of milestoning and temporal data modeling (unitemporal, bitemporal, etc.); Temporal Data Modeling — managing state changes over time (e.g., SCD Type 2).
- Strong experience with data reconciliation frameworks (using complex SQL queries) that perform various recon checks per industry standards to ensure migrated data is functionally equivalent.
- Positive understanding of Schema Evolution Management strategies and enforcement approaches.
Extraction Types Used
Kafka, ANSI SQL, FTP, Apache Spark
Data Formats
JSON, Avro, Parquet
Platforms
Hadoop (HDFS/Hive), Snowflake, Apache Iceberg, Sybase IQ
📌 Data Engineer — Datastore Migration (India)
🏢 Bounteous
📍 India