Detailed Job Description
Job Responsibilities:
- Responsible for ingesting data from multiple sources into a Data Lake using Java-based tools and frameworks such as Apache Spark, Spring Batch, and connectors for relational and NoSQL databases.
- Write transformation logic using Spark (Java API) and store processed data in HDFS or cloud storage solutions (e.g., AWS S3, GCP Storage, or on-prem equivalents).
- Responsible for writing stored procedures for relational databases (e.g., PostgreSQL, MySQL, or Oracle).
- Develop/re-develop existing ETL code and workflows into the modernized data pipeline.
- Experience with Batch and Realtime data processing and transformations.
Minimum Job Requirements:
- Must have experience in writing SQL queries and Spark code using Java.
- Must have experience with Java frameworks for data processing and integration (e.g., Spring, Apache Spark, JDBC).
- Should be familiar with Data Warehousing concepts and data modeling techniques like star schema.
Preferred Job Requirements:
- Positive to have knowledge of Delta Lake or similar transactional storage formats in Spark.
- Knowledge of various Slowly Changing Dimensions (SCD) types.
- Basics of workflow orchestration tools (e.g., Apache Airflow, Oozie) and trigger mechanisms.
- Strong client communication skills.
Educational Requirement - BE/BTech
Experience Level (Min-Max) - 5+ years
java,etl,sql,aws,data processing,
📌 Lead I (Pune)
🏢 UST
📍 Pune