28 Sep
|
Infosys
|
Bangalore East
28 Sep
Infosys
Bangalore East
- Primary skills: Python - Data Processing/Technology->Big Data - Data Processing->PySpark,Technology->OpenSystem->Python - OpenSystem->Python
Key Responsibilities
- Build and maintain Python-based data processing components for structured and semi-structured datasets.
- Develop and support ETL workflows to ingest, transform, validate, and load data into target systems.
- Design and optimize SQL queries for data extraction, transformation, reconciliation, and reporting needs.
- Implement reliable data pipelines with proper logging, error handling, retries, and monitoring hooks.
- Perform data quality checks, anomaly detection rules, and reconciliation to ensure accuracy and completeness.
- Troubleshoot pipeline failures and performance bottlenecks; drive root-cause analysis and permanent fixes.
- Collaborate with cross-functional teams to gather requirements and deliver incremental improvements.
- Maintain clear technical documentation for pipeline design, data mappings, and operational runbooks.
Minimum
Qualifications:
- Bachelor’s degree (or equivalent) in Engineering/Computer Science/IT or related field (BTech/BE/BSc or equivalent).
- 2–3 years of hands-on experience in Python for data processing and automation.
- Practical experience building ETL processes and working with data pipelines end-to-end.
- Strong SQL skills including joins, aggregations, subqueries, and performance-aware query writing.
- Ability to write clean, maintainable code and follow basic engineering practices (version control, reviews, testing mindset).
Preferred
Qualifications:
- Experience designing scalable pipeline patterns (incremental loads, CDC concepts, partitioning, backfills).
- Familiarity with Python data libraries and processing approaches (e.g., Pandas, batch processing patterns).
- Exposure to orchestration/scheduling concepts and operationalizing pipelines for reliability and observability.
- Experience working with large datasets and optimizing end-to-end pipeline performance (I/O, SQL tuning, compute efficiency).
- Proven ability to collaborate with stakeholders, translate requirements into technical solutions, and deliver within timelines.
Positive to have skills: Pandas, NumPy, Apache Airflow, Spark (PySpark), Linux/Shell Scripting
📌 Python - Data Processing (Bangalore East)
🏢 Infosys
📍 Bangalore East