- Understanding of Spark core concepts like RDD’s, DataFrames, DataSets, SparkSQL and Spark Streaming.
- Experience with Spark optimization techniques.
- Deep knowledge of Delta Lake features like time travel, schema evolution, data partitioning.
- Ability to design and implement data pipelines using Spark and Delta Lake as the data storage layer.
- Proficiency in Python/Scala/Java for Spark development and integrate with ETL process.
- Knowledge of data ingestion techniques from various sources (flat files, CSV, API, database)
- Understanding of data quality best practices and data validation techniques.
Other Skills:
- Understanding of data warehouse concepts, data modelling techniques.
- Expertise in Git for code management.
- Familiarity with CI/CD pipelines and containerization technologies.
- Nice to have experience using data integration tools like DataStage/Prophecy/Informatica/Ab Initio"
📌 Spark & Delta Lake (Hyderabad)
🏢 Capgemini
📍 Hyderabad
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.