- Join a team where data is treated as a product and engineering craftsmanship matters
- In this role you ll help build and optimize reliable data processing solutions that turn raw messy inputs into clean trusted datasets ready for analytics and downstream applications
- You ll collaborate closely with engineers analysts and stakeholders to understand data needs improve data quality and deliver pipelines that are scalable and easy to maintain
- If you enjoy solving real world data challenges writing clean Python code and taking ownership of end to end data workflows from ingestion to transformation and delivery this is a excellent opportunity to grow your impact
- You ll work in a supportive environment that values learning thoughtful problem solving and continuous improvement while delivering meaningful outcomes through data
Key Responsibilities:
- Key Responsibilities
- Design develop and maintain Python based data processing workflows for structured and semi structured datasets
- Build and optimize ETL processes to ingest transform validate and publish data for downstream consumption
- Develop and manage data pipelines ensuring reliability scalability and performance across workloads
- Write efficient SQL queries for data extraction transformation reconciliation and reporting needs
- Implement data quality checks validation rules and monitoring to ensure accuracy and completeness of datasets
- Troubleshoot pipeline failures identify root causes and implement preventive fixes to reduce recurrence
- Collaborate with cross functional teams to gather requirements define data contracts and deliver well documented solutions
- Maintain clear technical documentation for pipeline logic transformations and operational runbooks
- Follow best practices for code quality version control testing and peer reviews to ensure maintainable delivery
- Minimum Qualifications
- BTECH MTECH MCA or MSC in Computer Science Information Technology Data Software Engineering or a related field
- 3 5 years of experience in Python development focused on data processing and transformation use cases
- Hands on experience building ETL workflows and maintaining production grade data pipelines
- Strong SQL skills with experience in writing complex queries and optimizing performance
- Practical understanding of data processing concepts such as batching partitioning schema handling and data validation
Technical Requirements:
- PYTHON DATA PROCESSING DATAPIPELINE ETL Pandas NumPy Apache Airflow PySpark Data Quality Frameworks
Additional Responsibilities:
- Preferred Qualifications
- Experience designing scalable pipeline patterns incremental loads CDC style approaches idempotent processing and improving pipeline reliability
- Strong debugging skills with the ability to analyze data issues across multiple pipeline stages and resolve them efficiently
- Familiarity with orchestration and scheduling concepts for ETL jobs and dependency management
- Exposure to performance tuning for Python and SQL workloads including memory efficient processing and query optimization
- Proven ability to work with stakeholders to translate data requirements into robust well tested implementations
Preferred Skills:
Technology->OpenSystem->Python - OpenSystem->Python,Technology->Big Data - Data Processing->Spark
📌 Python - Data Processing (Bengaluru)
🏢 Infosys
📍 Bengaluru
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.