JUNIOR DATA ENGINEER —
1–3 years of hands-on experience with PySpark (DataFrames, Spark SQL basics) — production or solid academic/project experience acceptable.
Working knowledge of Fivetran or a similar ELT tool (Airbyte, Stitch) — setting up connectors, monitoring syncs.
Solid SQL fundamentals and exposure to a cloud data warehouse (Snowflake / Redshift / BigQuery).
Basic Python scripting ability for data tasks.
Exposure to at least one cloud platform (AWS / Azure / GCP), internship or project-level acceptable.
Eager to learn, works well under guidance from senior engineers; not expected to design architecture independently.
SENIOR DATA ENGINEER — REQUIREMENTS
5+ years of hands-on experience with PySpark (DataFrames, Spark SQL, RDDs, performance tuning, job optimization) in production.
3+ years configuring and managing Fivetran connectors at scale (custom connectors, schema drift handling, sync orchestration, troubleshooting).
Solid SQL and proven data modeling / warehouse design experience (Snowflake / Redshift / BigQuery).
Deep experience with at least one major cloud platform (AWS, Azure, or GCP), including cost and performance optimization.
Proven ability to design end-to-end pipeline architecture and lead technical decisions.
Experience mentoring junior engineers and reviewing code/pipeline designs.
RESPONSIBILITIES
Design, build, and optimize PySpark ETL/ELT pipelines for large-scale batch and/or streaming data.
Own Fivetran connector strategy across [X] source systems, including custom connector development.
Define data architecture and standards; review junior engineers' pipeline designs.
Collaborate with analytics/BI teams and business stakeholders to define data requirements.
Implement data quality frameworks, validation, and monitoring across pipelines.
📌 Data Engineer Pune
🏢 Sonata Software
📍 Pune
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.