We are seeking a highly skilled AWS Data Engineer with deep expertise in AWS
cloud architecture, big data processing, real-time streaming, and modern data
lake technologies. The ideal candidate will have solid hands-on experience in
Spark (PySpark), Iceberg, EMR, Starburst/Trino, and event-driven architectures,
along with experience building real-time and API-driven data applications who
can design and build generic solutions for one of our Fortune 500 Client
programs in the realm of Financial Master & Reference Data Management. This is
high visibility, fast-paced key initiative will integrate data across internal
and external sources, provide analytical insights, and integrate with the
customer’s critical systems.
RESPONSIBILITIES
Key Responsibilities
* Design and implement scalable, secure, and cost-optimized AWS data
architectures.
* Develop and maintain ETL pipelines using AWS Lambda and AWS Glue ETL.
* Configure and manage AWS Glue Crawlers, Glue Data Catalog, and schema
evolution.
* Build, optimize, and unit test applications on the Apache Spark framework
using PySpark.
* Design and optimize data lakes using Apache Iceberg on AWS, including table
compaction and Iceberg performance tuning.
* Work extensively with data formats such as Avro, Parquet, JSON, XML, and CSV.
* Orchestrate event-driven workflows using AWS Step Functions and Amazon
EventBridge.
* Connect and integrate Starburst from Lambda and Glue ETL jobs for federated
querying.
* Implement CI/CD pipelines for automated testing and deployment.
* Perform unit testing using PyTest, and performance tuning of Spark and Python
applications
QUALIFICATIONS
* Strong understanding of AWS architecture best practices, scalability,
security, and cost optimization strategies.
* Strong hands-on experience with AWS services including Lambda, Glue ETL,
Athena, S3, DynamoDB, Step Functions, EventBridge, SNS, and SQS.
* Deep experience in Apache Spark (PySpark/Scala) development, unit testing,
and performance opt
📌 AWS Data Engineer (Pune)
🏢 EXL
📍 Pune