Key Responsibilities
Build and maintain scalable data pipelines across stages such as data download, ingestion, and analysis within a bespoke data platform.
Develop high-performance data processing logic using Python and PySpark.
Perform large-scale transaction data analysis using custom algorithms and in-house logic to detect financial crime patterns.
Work with high-volume datasets and optimize pipeline performance.
Design productive data models and transformations for large-scale processing.
Write optimized SQL queries on PostgreSQL (RDS), leveraging: Window functions / Partitioning /Query performance optimization
Languages: Python, SQL
AWS exposure (especially EMR and S3)
Work with Delta tables and Parquet-based storage formats for efficient data processing.
Build and maintain Spark batch and streaming jobs.
Implement engineering best practices including unit testing, static code analysis, and CI/CD practices.
Contribute to data platform architecture and system design decisions.
Required Skills
Solid Python programming expertise
Deep understanding of: Functional programming/ Object-Oriented Programming (OOP)/ Design patterns/ Python execution and invocation mechanisms
Hands-on experience with Apache Spark / PySpark
Solid SQL expertise including: Window functions/ Partitioning/ Query optimization
Experience building end-to-end data pipelines
Data Platform Engineering
Experience designing scalable data processing architectures
Solid understanding of distributed data processing
Familiarity with standard data pipeline engineering practices
AWS exposure (especially EMR and S3)
Experience with Apache Airflow for job orchestration and scheduling
📌 Senior Aws Data Engineer Gurugram
🏢 Bounteous
📍 Gurugram
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.