Detailed Job Responsibilities: Responsible for ingesting data from multiple sources into a Data Lake using Java-based tools and frameworks such as Apache Spark, Spring Batch, and connectors for relational and NoSQL databases. Write transformation logic using Spark (Java API) and store processed data in HDFS or cloud storage solutions (e.g., AWS S3, GCP Storage, or on-prem equivalents). Responsible for writing stored procedures for relational databases (e.g., PostgreSQL, MySQL, or Oracle).
Develop/re-develop existing ETL code and workflows into the modernized data pipeline.
Experience with Batch and Realtime data processing and transformations.
Minimum Job Requirements: Must have experience in writing SQL queries and Spark code using Java.
Must have experience with Java frameworks for data processing and integration (e.g., Spring, Apache Spark, JDBC). Should be familiar with Data Warehousing concepts and data modeling techniques like star schema.
Preferred Job Requirements: Good to have knowledge of Delta Lake or similar transactional storage formats in Spark. Knowledge of various Slowly Changing Dimensions (SCD) types. Basics of workflow orchestration tools (e.g., Apache Airflow, Oozie) and trigger mechanisms. Solid client communication skills.
Educational
Requirement - BE/BTech Experience Level (Min-Max) - 5+ years
Skills
Java, Apache Spark, ETL, AWS, SQL
📌 Lead II - Data Engineering (Pune)
🏢 UST
📍 Pune
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.