- Build and maintain scalable data pipelines across stages such as data download, ingestion, and analysis within a bespoke data platform.
- Develop high-performance data processing logic using Python and PySpark.
- Perform large-scale transaction data analysis using custom algorithms and in-house logic to detect financial crime patterns.
- Work with high-volume datasets and optimize pipeline performance.
- Design efficient data models and transformations for large-scale processing.
- Write optimized SQL queries on PostgreSQL (RDS), leveraging: Window functions / Partitioning /Query performance optimization
- Languages: Python, SQL
- AWS exposure (especially EMR and S3)
- Work with Delta tables and Parquet-based storage formats for efficient data processing.
- Build and maintain Spark batch and streaming jobs.
- Implement engineering best practices including unit testing, static code analysis, and CI/CD practices.
- Contribute to data platform architecture and system design decisions.
Required Skills
- Strong Python programming expertise
- Deep understanding of: Functional programming/ Object-Oriented Programming (OOP)/ Design patterns/ Python execution and invocation mechanisms
- Hands-on experience with Apache Spark / PySpark
- Solid SQL expertise including: Window functions/ Partitioning/ Query optimization
- Experience building end-to-end data pipelines
- Data Platform Engineering
- Experience designing scalable data processing architectures
- Solid understanding of distributed data processing
- Familiarity with standard data pipeline engineering practices
- AWS exposure (especially EMR and S3)
- Experience with Apache Airflow for job orchestration and scheduling
📌 Senior AWS Data Engineer (Gurugram)
🏢 Bounteous
📍 Gurugram
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.