- Design, develop, test, deploy and maintain large-scale data pipelines using ETL tools such as PySpark and Spark Streaming.
- Collaborate with cross-functional teams to gather requirements for data processing needs and design scalable solutions.
- Develop complex data models using Delta Lake and Unity Catalog to support business intelligence initiatives.
- Troubleshoot issues related to Kafka messaging queues, AWS services, Microsoft Azure resources.
Job Requirements :
- 12-19 years of experience in Data Engineering with expertise in ETL (Extract Transform Load) processes.
- Robust proficiency in Python programming language with experience working with libraries like Pandas, NumPy, Matplotlib.
- Experience with big data technologies like Apache Spark, Hadoop ecosystem components including HDFS, YARN & MapReduce.