Pyspark Spark Bigdata Architect - Chennai/ Bangalore (Bengaluru)

Pyspark Spark Bigdata Architect - Chennai/ Bangalore (Bengaluru)

24 Sep
|
Tech Mahindra
|
Bengaluru

24 Sep

Tech Mahindra

Bengaluru

Primary Skill --

Python, Spark SQL, PySpark, Apache Iceberg Developer

Skill Bigdata Tech Architect

Exp – 12- 20 Yrs

Location – Chennai – Sholinganallur / Bangalore – RMZ Ecoworld

We are looking for experienced Data Engineers to join a large-scale data migration and modernization initiative. The role involves migrating data from enterprise relational databases to Apache Iceberg using PySpark, ensuring high-quality, performant, and reliable data pipelines.

The ideal candidate should possess strong SQL expertise, hands-on PySpark experience, good debugging and troubleshooting capabilities, and the ability to analyze and resolve production issues independently.

Key Responsibilities

- Analyze source system data models, schemas, and business requirements to create effective migration strategies.
- Design, develop, and maintain scalable PySpark-based data migration pipelines for moving data from source to Apache Iceberg.
- Perform data extraction, transformation, cleansing, standardization, and loading activities while ensuring data integrity.
- Develop and optimize complex SQL queries for data extraction, validation, reconciliation, and reporting.
- Perform SQL query tuning and performance optimization to improve execution efficiency on large datasets.
- Conduct detailed data profiling and identify data quality issues prior to migration.
- Define and execute data validation and reconciliation processes to ensure completeness and accuracy of migrated data.
- Debug and troubleshoot PySpark jobs, identifying root causes for failures, performance bottlenecks, and data inconsistencies.




- Analyze Spark execution plans and optimize jobs through partitioning, caching, and efficient transformation techniques.
- Monitor and troubleshoot Azure DevOps (ADO) pipeline executions, resolve failures, and coordinate deployments.
- Work with Git repositories to manage code changes through commits, pull requests, code reviews, branching, and merge strategies.
- Investigate production issues, perform Root Cause Analysis (RCA), and implement preventive measures.
- Review job logs, application logs, and execution metrics to identify and resolve operational issues.
- Collaborate with business users, architects, data engineers, QA teams, and project stakeholders to ensure successful project delivery.
- Create and maintain technical documentation, migration runbooks, mapping documents, and operational procedures.
- Support SIT, UAT, deployment, hypercare, and production support activities.
- Ensure adherence to coding standards, development best practices, and data governance policies.
- Participate in estimation, planning, sprint activities, and technical discussions.

Good to Have
- Exposure to Apache Iceberg, Data Lake, or Lakehouse architectures
- Understanding of DAG/workflow analysis and pipeline dependency troubleshooting
- Ability to perform Root Cause Analysis (RCA) and resolve pipeline failures
- Experience with CI/CD processes and release management
- Knowledge of Spark performance tuning and monitoring
- Experience in production support and incident management
- Agile/Scrum project experience
- Robust analytical, troubleshooting, and problem-solving skills

📌 Pyspark Spark Bigdata Architect - Chennai/ Bangalore (Bengaluru)
🏢 Tech Mahindra
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: pyspark spark bigdata architect - chennai/ bangalore (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: pyspark spark bigdata architect - chennai/ bangalore (bengaluru) / bengaluru