24 Sep
|
Tech Mahindra
|
Bengaluru
24 Sep
Tech Mahindra
Bengaluru
Primary Skill --
Python, Spark SQL, PySpark, Apache Iceberg Developer
Skill Bigdata Tech Architect
Exp – 12- 20 Yrs
Location – Chennai – Sholinganallur / Bangalore – RMZ Ecoworld
We are looking for experienced Data Engineers to join a large-scale data migration and modernization initiative. The role involves migrating data from enterprise relational databases to Apache Iceberg using PySpark, ensuring high-quality, performant, and reliable data pipelines.
The ideal candidate should possess strong SQL expertise, hands-on PySpark experience, good debugging and troubleshooting capabilities, and the ability to analyze and resolve production issues independently.
Key Responsibilities
- Analyze source system data models, schemas, and business requirements to create effective migration strategies.
- Design, develop, and maintain scalable PySpark-based data migration pipelines for moving data from source to Apache Iceberg.
- Perform data extraction, transformation, cleansing, standardization, and loading activities while ensuring data integrity.
- Develop and optimize complex SQL queries for data extraction, validation, reconciliation, and reporting.
- Perform SQL query tuning and performance optimization to improve execution efficiency on large datasets.
- Conduct detailed data profiling and identify data quality issues prior to migration.
- Define and execute data validation and reconciliation processes to ensure completeness and accuracy of migrated data.
- Debug and troubleshoot PySpark jobs, identifying root causes for failures, performance bottlenecks, and data inconsistencies.
- Analyze Spark execution plans and optimize jobs through partitioning, caching, and efficient transformation techniques.
- Monitor and troubleshoot Azure DevOps (ADO) pipeline executions, resolve failures, and coordinate deployments.
- Work with Git repositories to manage code changes through commits, pull requests, code reviews, branching, and merge strategies.
- Investigate production issues, perform Root Cause Analysis (RCA), and implement preventive measures.
- Review job logs, application logs, and execution metrics to identify and resolve operational issues.
- Collaborate with business users, architects, data engineers, QA teams, and project stakeholders to ensure successful project delivery.
- Create and maintain technical documentation, migration runbooks, mapping documents, and operational procedures.
- Support SIT, UAT, deployment, hypercare, and production support activities.
- Ensure adherence to coding standards, development best practices, and data governance policies.
- Participate in estimation, planning, sprint activities, and technical discussions.
Good to Have
- Exposure to Apache Iceberg, Data Lake, or Lakehouse architectures
- Understanding of DAG/workflow analysis and pipeline dependency troubleshooting
- Ability to perform Root Cause Analysis (RCA) and resolve pipeline failures
- Experience with CI/CD processes and release management
- Knowledge of Spark performance tuning and monitoring
- Experience in production support and incident management
- Agile/Scrum project experience
- Robust analytical, troubleshooting, and problem-solving skills
📌 Pyspark Spark Bigdata Architect - Chennai/ Bangalore (Bengaluru)
🏢 Tech Mahindra
📍 Bengaluru