Role & responsibilities : Pyspark ETL engineer
Position: Senior PySpark ETL Engineer
Role Summary
The Senior PySpark ETL Engineer is responsible for designing, building, optimizing, and
operating scalable data pipelines using Apache Spark (PySpark). This role focuses on
highvolume batch (and optionally streaming) data processing, ensuring
performance, reliability, data quality, and cost efficiency across enterprise data
platforms.
The position requires strong python, hands on Spark expertise, deep SQL and data
modeling knowledge, and the ability to own pipelines end toend in production.
Mandatory Requirements
1012 years of overall IT experience, with strong focus on data engineering
and ETL.
3+ years of hands on experience with PySpark / Apache Spark in production
environments.
Strong experience designing and implementing ETL / ELT pipelines at scale.
Excellent knowledge of SQL and relational data concepts.
Experience handling large datasets in distributed environments.
Solid ownership mindset, problemsolving skills, and ability to independently
handle production pipelines.
Core Technical Skills
PySpark & Spark Engineering
Deep expertise in PySpark:
DataFrames, Spark SQL, window functions, joins, aggregations
Spark execution model (DAGs, stages, tasks)
Strong hands on experience with:
Partitioning strategies
Shuffle optimization
Broadcast vs sortmerge joins
Caching / persisting
Handling data skew and memory spills
Proven ability to debug and optimize slow Spark jobs.
ETL & Data Engineering
Strong knowledge of ETL/ELT design patterns:
Incremental loads
Watermarking
Idempotent pipeline design
Reprocessing and backfill strategies
Experience implementing:
SCD Type 1 / Type 2
Deduplication and late arriving data handling
Ability to design reusable transformation frameworks and common utilities.
Experience building source to target reconciliation and data quality checks.
Data Storage & SQL
Excellent SQL skills including:
Complex joins
Subqueries and CTEs
Window functions
Query optimization
Experience working with:
RDBMS sources (Postgres, MySQL)
Data lake storage using Parquet / ORC
Experience with partitioned datasets and compaction strategies.
Cloud & Big Data Platforms
Hands on experience with at least one Spark platform:
AWS EMR
Spark on Kubernetes
Experience working with cloud storage:
S3
Familiarity with orchestration tools:
Airflow, Databricks Workflows, ADF, or equivalent.
Responsibilities
Pipeline Development & Ownership
Design, implement, and maintain highperformance PySpark ETL pipelines.
Own pipelines end toend, including development, deployment, monitoring, and
production support.
Ensure pipelines are scalable, fault tolerant, and re runnable.
Implement incremental processing and efficient data movement strategies.
Performance & Reliability
Identify and fix Spark performance bottlenecks.
Optimize resource usage and reduce execution time and cost.
Handle production issues related to:
Job failures
Data corruption
SLA breaches
Perform root cause analysis and implement permanent fixes.
Data Quality & Governance
Implement strong data quality validations, checks, and reconciliation
mechanisms.
Ensure correctness, completeness, and freshness of datasets.
Follow enterprise standards for:
Data retention
Auditability
Schema evolution
Engineering Excellence
Write clean, maintainable, and testable PySpark code.
Conduct code reviews and guide junior engineers.
Follow best practices for:
Version control (Git)
CI/CD
Logging and monitoring
Maintain clear documentation and operational runbooks
If any candidate interested please connect in
[email protected]
Preferred candidate profile
📌 Pyspark Developer (Chennai)
🏢 SightSpectrum
📍 Chennai