We are seeking an enthusiastic, detail-oriented ETL Developer to join our data engineering team. In this role, you will be designing, developing, and maintaining scalable data pipelines using PySpark, Talend, and Ab Initio. You will work closely with senior developers to transform raw data into valuable business insights, ensuring optimal data quality and performance
ETL Pipeline Development: Building and maintaining Extraction, Transformation, and Loading (ETL) workflows
Big Data Processing: Utilize PySpark for data processing, manipulation, and analysis within big data ecosystems.
Data Extraction & Ingestion: Extract data from multiple sources including relational databases (SQL), flat files etc.
Data Transformation: Create mapping documents and apply business logic to transform large datasets.
Performance Tuning: Assist in analyzing and optimizing existing ETL jobs and Spark scripts for better scalability and execution speed.
Testing & Documentation: Perform unit testing, validate data accuracy, and maintain technical documentation for all ETL processes.
Production Support:
Help monitor scheduled daily/weekly jobs and troubleshoot data integration or quality issues
Recommended Qualifications:
8-10 years of relevant experience
Experience in systems analysis and programming of software applications
Programming Languages: Proficiency in Pyspark and SQL.
Big Data Frameworks: Hands-on experience or robust conceptual understanding of PySpark.
ETL Tools: Familiarity with Data Integration/Big Data tools and Ab Initio components.
Scripting: Basic to intermediate knowledge of Unix/Linux Shell Scripting.
Databases: Understanding of Big Data (Hadoop) & RDBMS like Oracle, or PostgreSQL
Ability to work under pressure and manage deadlines or unexpected changes in expectations or requirements
Education:
Bachelor’s degree/University degree or equivalent experience
This job description provides a high-level review of the types of work performed. Other job-related duties may be a
📌 Applications Developer (India)
🏢 Citi
📍 India