29 Aug
|
Talentgigs
|
Chennai
29 Aug
Talentgigs
Chennai
Position: Senior PySpark ETL Engineer
About the Role The Senior PySpark ETL Engineer is responsible for designing, building, optimizing, and operating scalable data pipelines using Apache Spark (PySpark). This role focuses on high-volume batch (and optionally streaming) data processing, ensuring performance, reliability, data quality, and cost efficiency across enterprise data platforms. The position requires strong python, hands-on Spark expertise, deep SQL and data modeling knowledge, and the ability to own pipelines end-to-end in production.
Responsibilities
- Design, implement, and maintain high-performance PySpark ETL pipelines.
- Own pipelines end-to-end, including development, deployment, monitoring, and production support.
- Ensure pipelines are scalable, fault tolerant, and re-runnable.
- Implement incremental processing and efficient data movement strategies.
- Identify and fix Spark performance bottlenecks.
- Optimize resource usage and reduce execution time and cost.
- Handle production issues related to job failures, data corruption, and SLA breaches.
- Perform root cause analysis and implement permanent fixes.
- Implement solid data quality validations, checks, and reconciliation mechanisms.
- Ensure correctness, completeness, and freshness of datasets.
- Follow enterprise standards for data retention, auditability, and schema evolution.
- Write clean, maintainable, and testable PySpark code.
- Conduct code reviews and guide junior engineers.
- Follow best practices for version control (Git), CI/CD, logging, and monitoring.
- Maintain clear documentation and operational runbooks.
Qualifications
- 10–12 years of overall IT experience,
with strong focus on data engineering and ETL.
- 3+ years of hands-on experience with PySpark / Apache Spark in production environments.
- Strong experience designing and implementing ETL / ELT pipelines at scale.
- Excellent knowledge of SQL and relational data concepts.
- Experience handling large datasets in distributed environments.
- Strong ownership mindset, problem-solving skills, and ability to independently handle production pipelines.
Required Skills
- Deep expertise in PySpark: DataFrames, Spark SQL, window functions, joins, aggregations.
- Strong hands-on experience with partitioning strategies, shuffle optimization, broadcast vs sort-merge joins, caching/persisting, handling data skew and memory spills.
- Proven ability to debug and optimize slow Spark jobs.
Preferred Skills
- Strong knowledge of ETL/ELT design patterns: Incremental loads, watermarking, idempotent pipeline design, reprocessing and backfill strategies.
- Experience implementing SCD Type 1 / Type 2, deduplication and late arriving data handling.
- Ability to design reusable transformation frameworks and common utilities.
- Experience building source to target reconciliation and data quality checks.
- Excellent SQL skills including complex joins, subqueries and CTEs, window functions, and query optimization.
- Experience working with RDBMS sources (Postgres, MySQL) and data lake storage using Parquet / ORC.
- Experience with partitioned datasets and compaction strategies.
- Hands-on experience with at least one Spark platform: AWS EMR, Spark on Kubernetes.
- Experience working with cloud storage: S3.
- Familiarity with orchestration tools: Airflow, Databricks Workflows, ADF, or equivalent.
📌 Senior PySpark ETL Engineer (Chennai)
🏢 Talentgigs
📍 Chennai