- Design and develop scalable ETL and data processing pipelines on AWS.
- Build batch and real-time data pipelines using PySpark, Spark Streaming, and Kafka.
- Develop and optimize Data Lake solutions using Apache Hudi or Apache Iceberg.
- Build and manage data processing workflows using AWS Glue and Amazon EMR.
- Automate infrastructure deployment using Terraform and containerize applications using Docker and Amazon ECS.
- Optimize large-scale distributed data processing for performance, scalability, and reliability.
- Implement data cataloging and metadata management using AWS Glue Data Catalog.
- Collaborate with cross-functional teams to deliver secure, high-quality data engineering solutions.
- Troubleshoot production issues and continuously improve data platform performance.
Key Skills
- 8+ years of experience in Data Engineering.
- Robust expertise in AWS Glue, PySpark, Amazon EMR, and Spark Streaming.
- Hands-on experience with Apache Kafka for real-time data streaming.
- Experience with Apache Hudi or Apache Iceberg.
- Strong programming skills in Python and Object-Oriented Programming (OOPS).
- Experience with AWS Glue ETL and Glue Data Catalog.
- Hands-on experience with Docker, Amazon ECS, and Terraform.
📌 Data Engineer (Pune)
🏢 Tekpillar
📍 Pune
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.