21 Aug
|
Relevantz Technology Services
|
Chennai
21 Aug
Relevantz Technology Services
Chennai
Position: Senior PySpark ETL Engineer
Role Summary The Senior PySpark ETL Engineer is responsible for designing, building, optimizing, and operating scalable data pipelines using Apache Spark (PySpark). This role focuses on high volume batch (and optionally streaming) data processing, ensuring performance, reliability, data quality, and cost eciency across enterprise data platforms.
The position requires robust python, hands on Spark expertise, deep SQL and data modeling knowledge, and the ability to own pipelines end to end in production.
Mandatory Requirements
10 - 12 years of overall IT experience, with strong focus on data engineering and ETL.
3+ years of hands on experience with PySpark / Apache Spark in production environments.
Strong experience designing and implementing ETL / ELT pipelines at scale.
Excellent knowledge of SQL and relational data concepts.
Experience handling large datasets in distributed environments.
Robust ownership mindset, problem solving skills, and ability to independently handle production pipelines.
Core Technical Skills
PySpark & Spark Engineering
Deep expertise in PySpark
DataFrames, Spark SQL, window functions, joins, aggregations
Spark execution model (DAGs, stages, tasks)
Strong hands on experience with:
Partitioning strategies
Shuffle optimization
Broadcast vs sort merge joins
Caching / persisting
Handling data skew and memory spills
Proven ability to debug and optimize slow Spark jobs.
ETL & Data Engineering:
Strong knowledge of ETL/ELT design patterns:
Incremental loads
Watermarking
Idempotent pipeline design
Reprocessing and backfill strategies
Experience implementing
SCD Type 1 / Type 2
Deduplication and late arriving data handling
Ability to design reusable transformation frameworks and common utilities.
Experience building source to target reconciliation and data quality checks.
Data Storage & SQL
Excellent SQL skills including
Complex joins
Subqueries and CTEs
Window functions
Query optimization
Experience working with
RDBMS sources (Postgres, MySQL)
Data lake storage using Parquet / ORC
Experience with partitioned datasets and compaction strategies.
Cloud & Big Data Platforms
Hands on experience with at least one Spark platform:
AWS EMR
Spark on Kubernetes
Experience working with cloud storage:
S3
Familiarity with orchestration tools
Airflow, Databricks Workflows, ADF, or equivalent.
Responsibilities
Pipeline Development & Ownership
Design, implement, and maintain high performance PySpark ETL pipelines.
Own pipelines end to end, including development, deployment, monitoring, and production support.
Ensure pipelines are scalable, fault tolerant, and re runnable.
Implement incremental processing and ecient data movement strategies.
Performance & Reliability
Identify and fix Spark performance bottlenecks.
Optimize resource usage and reduce execution time and cost.
Handle production issues related to:
Job failures
Data corruption
SLA breaches
Perform root cause analysis and implement permanent fixes.
Data Quality & Governance
Implement strong data quality validations, checks, and reconciliation mechanisms.
Ensure correctness, completeness, and freshness of datasets.
Follow enterprise standards for
Data retention
Auditability
Schema evolution
Engineering Excellence
Write clean, maintainable, and testable PySpark code.
Conduct code reviews and guide junior engineers.
Follow best practices for
Version control (Git)
CI/CD
Logging and monitoring
Maintain clear documentation and operational runbooks.
📌 Senior PySpark ETL Lead Engineer (Chennai)
🏢 Relevantz Technology Services
📍 Chennai