Requirements
- Strong understanding of AWS architecture best practices, scalability, security, and cost optimization strategies.
- Strong hands-on experience with AWS services including Lambda, Glue ETL, Athena, S3, DynamoDB, Step Functions, EventBridge, SNS, and SQS.
- Deep experience in Apache Spark (PySpark/Scala) development, unit testing, and performance optimization.
- Robust Python programming skills using libraries such as pandas, requests, json, and awswrangler.
- Experience on Apache Kafka and Confluent Kafka.
- Experience designing and optimizing data lakes using Apache Iceberg, including compaction and Iceberg optimization techniques.
- Hands-on experience integrating and optimizing Starburst/Trino, including connecting Starburst from Lambda and Glue ETL jobs.
- Experience with NoSQL databases such as DynamoDB, MongoDB.
- Experience working with data formats including Avro, Parquet, JSON, XML, and CSV.
- Comfortable challenging your peers and leadership team.
- Can prove yourself quickly and decisively.
- Excellent communication skills and Good Customer Centricity
Design and optimize data lakes using Apache Iceberg on AWS, including table compaction and Iceberg performance tuning.
📌 Data Engineer (Pune)
🏢 EXL
📍 Pune