Position Summary
We are looking for a highly experienced Senior ETL/Data Engineer with strong expertise in Talend Data Integration, PySpark, Oracle, and up-to-date cloud-based data engineering practices. The ideal candidate will have extensive experience designing, analyzing, enhancing, and modernizing enterprise ETL solutions while working closely with business and technical stakeholders. This role requires a strong understanding of data warehousing, ETL architecture, legacy system analysis, and cloud-based data transformation frameworks.
The candidate should be comfortable reverse engineering existing ETL processes, implementing scalable data pipelines, and supporting enterprise data migration initiatives.
Mandatory Skills & Experience
- 10+ years of experience in ETL/Data Engineering.
- Strong hands-on experience with Talend Open Studio / Talend Data Integration.
- Strong expertise in PySpark and cloud-based data engineering technologies.
- Experience reverse engineering existing Talend ETL jobs and understanding complex ETL workflows.
- Strong Java fundamentals (generated Talend code understanding).
- Strong SQL and PL/SQL programming skills.
- Hands-on experience with Oracle Database (Oracle 19c preferred).
- Strong understanding of Data Warehousing concepts including:
- Star Schema
- Snowflake Schema
- Facts & Dimensions
- Experience designing enterprise ETL architectures.
- Strong experience with
- Data Transformation
- Incremental Loading
- Change Data Capture (CDC)
- Batch Processing
- Fact Table Loading Strategies
- Experience understanding and translating complex business rules into ETL logic.
- Hands-on experience with Jenkins for ETL job orchestration and deployment.
- Experience with Git or similar Version Control Systems.
- Strong debugging, troubleshooting, analytical, and problem-solving skills.
- Experience preparing technical documentation and Source-to-Target Mapping (STM).
- Excellent verbal and written communication skills with the ability to interact directly with client stakeholders.
Good-to-Have Skills
- Working knowledge of AWS services including:
- Amazon EC2
- Amazon S3
- Amazon WorkSpaces
- Experience with modern cloud data platforms such as:
- Databricks
- AWS Glue
- PySpark
- Experience in ETL modernization and legacy platform migration initiatives.
- Oracle SQL and ETL performance tuning.
- Data Modeling experience.
- Exposure to CI/CD best practices in Data Engineering.
- Knowledge of enterprise data governance and data quality practices.
Preferred Candidate Profile The ideal candidate should possess:
- Strong ownership mindset with the ability to work independently.
- Excellent analytical and logical reasoning skills.
- Ability to quickly understand complex legacy ETL ecosystems.
- Passion for modernizing traditional ETL platforms using cloud-native technologies.
- Strong stakeholder management and communication skills.
- Experience working in Agile/Scrum environments.
- Ability to mentor junior engineers and contribute to technical best practices.
Key Responsibilities
- Design, develop, enhance, and maintain enterprise-scale ETL solutions using Talend.
- Reverse engineer and analyze existing Talend jobs to understand business logic and data transformation processes.
- Develop scalable ETL pipelines using Talend, PySpark, SQL, and Oracle technologies.
- Build efficient data ingestion, transformation, validation, and loading processes.
- Implement incremental loads, Change Data Capture (CDC), and various ETL loading strategies.
- Design and optimize Fact and Dimension loading processes for enterprise Data Warehouses.
- Collaborate with business analysts and stakeholders to understand complex business rules and convert them into scalable ETL solutions.
- Optimize ETL performance, SQL queries, and Oracle database operations.
- Participate in migration and modernization initiatives involving legacy ETL platforms.
- Implement CI/CD practices using Jenkins and manage source code through Git.
- Prepare and maintain technical documentation including:
- Source-to-Target Mapping (STM)
- Technical Design Documents
- Data Flow Diagrams
- ETL Design Specifications
- Troubleshoot production issues, perform root cause analysis, and implement long-term fixes.
- Collaborate effectively with cross-functional teams including Architects, DBAs, Business Analysts, QA teams, and Client stakeholders.
📌 Data Engineer (Pune)
🏢 Cybage
📍 Pune