Data Engineer to build and maintain scalable data pipelines on AWS using PySpark, SQL, and Python, supporting analytics use cases in the insurance domain.
Key Responsibilities:
- • Build and maintain ETL/ELT pipelines (PySpark, SQL, Python)
- • Develop batch & near real-time ingestion pipelines
- • Implement incremental loads (CDC)
- • Work with AWS: S3, Glue, EMR, Redshift, Athena
- • Develop Airflow DAGs for orchestration
- • Support CI/CD pipelines (CodeCommit/Bitbucket)
- • Ensure data quality, validation, and monitoring
- • Optimize Spark jobs and SQL queries
- • Follow data governance and security practices
- • Collaborate with business and analytics teams