Key Responsibilities
- Design and develop scalable ETL and data processing pipelines on AWS.
- Build batch and real-time data pipelines using
PySpark, Spark Streaming, and Kafka
.
- Develop and optimize Data Lake solutions using
Apache Hudi
or
Apache Iceberg
.
- Build and manage data processing workflows using
AWS Glue
and
Amazon EMR
.
- Automate infrastructure deployment using
Terraform
and containerize applications using
Docker
and
Amazon ECS
.
- Optimize large-scale distributed data processing for performance, scalability, and reliability.
- Implement data cataloging and metadata management using
AWS Glue Data Catalog
.
- Collaborate with cross-functional teams to deliver secure, high-quality data engineering solutions.
- Troubleshoot production issues and continuously improve data platform performance.
Key Skills
- 8+ years of experience in Data Engineering.
- Strong expertise in
AWS Glue
,
PySpark
,
Amazon EMR
, and
Spark Streaming
.
- Hands-on experience with
Apache Kafka
for real-time data streaming.
- Experience with
Apache Hudi
or
Apache Iceberg
.
- Robust programming skills in
Python
and Object-Oriented Programming (OOPS).
- Experience with
AWS Glue ETL
and
Glue Data Catalog
.
- Hands-on experience with
Docker
,
Amazon ECS
, and
Terraform
.
📌 Data Engineer (Pune Division)
🏢 TekPillar®
📍 Pune Division
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.