18 Aug
|
Synthlane
|
India
This is a remote position.
Role Overview
We are looking for Data Engineers at Senior and Mid-Level to join our team in building a privacy-preserving data platform where data engineering meets production-grade software engineering.
You will work on designing, developing, and maintaining reliable data pipelines that transform operational data into high-quality, secure, and AI-ready datasets.
Key Responsibilities
- Build and maintain production-grade data pipelines on AWS.
- Extract, transform, validate, and curate large-scale Parquet datasets.
- Implement data de-identification, masking, and privacy-preserving transformations.
- Design and maintain data pipeline orchestration, scheduling, retries, and backfill mechanisms.
- Implement comprehensive data quality checks, monitoring, and alerting.
- Work with workflow orchestration tools such as Airflow, Dagster, or AWS Step Functions.
- Contribute to CI/CD pipelines and Infrastructure as Code (IaC) practices.
- Manage schema evolution and schema drift across data sources and pipelines.
- Provide production support, troubleshooting, and root cause analysis for data pipeline issues.
- Maintain data catalogs, metadata, and data lineage.
- Follow software engineering best practices including Git, code reviews, automated testing, and maintainable code.
- Build reliable and idempotent data pipelines capable of handling retries and large-scale backfills.
Required Skills & Experience
- Strong proficiency in Python and SQL.
- Hands-on experience with AWS data services and production data pipelines.
- Experience with Apache Spark or equivalent distributed data processing technologies.
- Practical experience with Airflow, Dagster, AWS Step Functions, or similar orchestration tools.
- Strong understanding of data pipeline architecture, ETL/ELT, and data transformation.
- Experience working with Parquet and large-scale datasets.
- Understanding of data quality, schema management, monitoring, and alerting.
- Solid software engineering practices including:
- Git and version control
- Code reviews
- Automated testing
- Idempotency
- Error handling
- Retries and backfills
- Experience supporting and troubleshooting production data pipelines.
- Ability to work effectively with cross-functional engineering and data teams.
Requirements
Required Skills
Python | SQL | AWS | Data Engineering | Data Pipelines | ETL/ELT | Apache Spark | Parquet | Airflow | Dagster | AWS Step Functions | Data Orchestration | Data Quality | Schema Management | Data Transformation | Production Support | Git | CI/CD | Automated Testing | Data De-identification | Data Lineage | Data Catalog
Good to Have
Debezium | AWS DMS | Apache Iceberg | Delta Lake | Apache Hudi | Data Masking | Data Tokenization | Terraform | CloudFormation | ML/AI Training Data | Privacy-Preserving Data
📌 Data Engineer – Mid & Senior Level (India)
🏢 Synthlane
📍 India