Requirements
n
n· Strong understanding of AWS architecture best practices, scalability, security, and cost optimization strategies.
n· Strong hands-on experience with AWS services including Lambda, Glue ETL, Athena, S3, DynamoDB, Step Functions, EventBridge, SNS, and SQS.
n· Deep experience in Apache Spark (PySpark/Scala) development, unit testing, and performance optimization.
n· Strong Python programming skills using libraries such as pandas, requests, json, and awswrangler.
n· Experience on Apache Kafka and Confluent Kafka.
n· Experience designing and optimizing data lakes using Apache Iceberg, including compaction and Iceberg optimization techniques.
n· Hands-on experience integrating and optimizing Starburst/Trino,
including connecting Starburst from Lambda and Glue ETL jobs.
n· Experience with NoSQL databases such as DynamoDB, MongoDB.
n· Experience working with data formats including Avro, Parquet, JSON, XML, and CSV.
n· Comfortable challenging your peers and leadership team.
n· Can prove yourself quickly and decisively.
n· Excellent communication skills and Valuable Customer Centricity
nDesign and optimize data lakes using Apache Iceberg on AWS, including table compaction and Iceberg performance tuning.
📌 Data Engineer (Pune)
🏢 EXL
📍 Pune