Discover your future at Citi:
Working at Citi is far more than just a job. A career with us means joining a team of more than 230,000 dedicated people from around the globe. At Citi, you’ll have the prospect to grow your career, give back to your community and make a real impact.
Job Overview:
You will work under the guidance of senior engineers and collaborate with data teams to build reliable data pipelines and contribute to analytics and reporting solutions.
Key Responsibilities
Development & Engineering
Assist in developing and maintaining data pipelines using Python and PySpark
Support ETL/ELT workflows for batch data processing
Write clean, readable, and well‑structured Python code following best practices
Perform basic data transformations, aggregations, and validations
Debug and troubleshoot pipeline issues with guidance from senior developers
Data & Platform
Work with structured and semi‑structured data formats (CSV, JSON, Parquet, etc.)
Assist in integrating data from databases, APIs, and cloud storage systems
Help ensure data quality and consistency within pipelines
Support migration of legacy scripts to up-to-date data platforms
Learning & Collaboration
Collaborate with team members on development tasks and code reviews
Participate in knowledge‑sharing and training sessions
Learn and adopt current tools, frameworks, and best practices
Assist in documenting data workflows and technical processes
Required Skills & Qualifications
Technical Skills
Basic to intermediate proficiency in Python
4 -7 years of experience
Exposure to Apache Spark / PySpark (internship or project experience is acceptable)
Understanding of fundamental programming and data structures
Basic knowledge of SQL and relational databases
Familiarity with data processing concepts and ETL fundamentals
Awareness of Linux/Unix command line is a plus
Engineering Fundamentals
Understanding of coding best practices and version control (Git)
Basic debugging
📌 Python And Pyspark Developer Chennai (India)
🏢 Citi
📍 India