Role & responsibilities : Pyspark ETL engineer
Position: Senior PySpark ETL Engineer
Role Summary The Senior PySpark ETL Engineer is responsible for designing, building, optimizing, and operating scalable data pipelines using Apache Spark (PySpark). This role focuses on highvolume batch (and optionally streaming) data processing, ensuring performance, reliability, data quality, and cost efficiency across enterprise data platforms.
The position requires strong python, hands on Spark expertise, deep SQL and data modeling knowledge, and the ability to own pipelines end toend in production.
Mandatory Requirements
1012 years of overall IT experience, with strong focus on data engineering and ETL.
3+ years of hands on experience with PySpark / Apache Spark in production environments.
Strong experience designing and implementing ETL / ELT pipelines at scale.
Excellent knowledge of SQL and relational data concepts.
Experience handling large datasets in distributed environments.
Robust ownership mindset, problemsolving skills, and ability to independently handle production pipelines.
Core Technical Skills
PySpark & Spark Engineering
Deep expertise in PySpark
DataFrames, Spark SQL, window functions, joins, aggregations
Spark execution model (DAGs, stages, tasks)
Strong hands on experience with:
Partitioning strategies
Shuffle optimization
Broadcast vs sortmerge joins
Caching / persisting
Handling data skew and memory spills
Proven ability to debug and optimize slow Spark jobs.
ETL & Data Engineering
Strong knowledge of ETL/ELT design patterns:
Incremental loads
Watermarking
Idempotent pipeline design
Reprocessing and backfill strategies
Experience implementing
SCD Type 1 / Type 2
Deduplication and late arriving data handling
Ability to design reusable transformation frameworks and common utilities.
Experience building source to target reconciliation and data quality checks.
Data Storage & SQL
Excellent SQL skills including
Complex joins
Subqueries and CTEs
Window functions
Query optimization
Experience working with
RDBMS sources (Postgres, MySQL)
Data lake storage using Parquet / ORC
Experience with partitioned datasets and compaction strategies.
Cloud & Big Data Platforms
Hands on experience with at least one Spark platform:
AWS EMR
Spark on Kubernetes
Experience working with cloud storage:
S3
Familiarity with orchestration tools
Airflow, Databricks Workflows, ADF, or equivalent.
Responsibilities
Pipeline Development & Ownership
Design, implement, and maintain highperformance PySpark ETL pipelines.
Own pipelines end toend, including development, deployment, monitoring, and production support.
Ensure pipelines are scalable, fault tolerant, and re runnable.
Implement incremental processing and efficient data movement strategies.
Performance & Reliability
Identify and fix Spark performance bottlenecks.
Optimize resource usage and reduce execution time and cost.
Handle production issues related to:
Job failures
Data corruption
SLA breaches
Perform root cause analysis and implement permanent fixes.
Data Quality & Governance
Implement strong data quality validations, checks, and reconciliation mechanisms.
Ensure correctness, completeness, and freshness of datasets.
Follow enterprise standards for
Data retention
Auditability
Schema evolution
Engineering Excellence
Write clean, maintainable, and testable PySpark code.
Conduct code reviews and guide junior engineers.
Follow best practices for
Version control (Git)
CI/CD
Logging and monitoring
Maintain clear documentation and operational runbooks
If any candidate interested please connect in
[email protected]
Preferred candidate profile
📌 Pyspark Developer (Chennai)
🏢 SightSpectrum
📍 Chennai