14 Aug
|
ClanX
|
Gurugram
Requirements
- 3–7 years of hands-on experience building production-grade data pipelines
- Strong expertise in SQL, including query optimization and complex transformations
- Strong Python programming skills for data engineering workloads
- Experience with Spark/PySpark or equivalent distributed processing frameworks
- Strong understanding of data modeling, ETL/ELT, CDC, incremental processing, and partitioning
- Experience with workflow orchestration tools such as Dagster, Airflow, or Prefect
- Experience working with AWS or Azure cloud platforms
- Strong focus on data quality, observability, validation, and monitoring
- Experience building scalable datasets for analytics and machine learning use cases
- Bachelor's degree in Computer Science, Engineering, Mathematics, or a related field
Responsibilities
- Design, build, and maintain scalable batch ETL/ELT and CDC pipelines
- Own data ingestion from client systems into the data platform
- Build reliable, idempotent,
incremental, and backfill-protected data pipelines
- Model and transform raw data into trusted analytical and ML-ready datasets
- Develop and optimize Spark/PySpark and SQL workloads for large-scale processing
- Implement data quality checks, lineage, monitoring, and alerting mechanisms
- Support machine learning feature pipelines and analytics platforms
- Optimize storage formats, partitioning strategies, and query performance
- Collaborate with ML, product, and engineering teams to deliver reliable data solutions
- Drive best practices for scalability, reliability, and maintainability across the data stack
Job Details
Gurugram - Hybrid (2–3 days on-site)
Interview Process
- Technical Round (Python & SQL)
- System Design Round
- Culture Fit Round
📌 Senior Data Engineer (Gurugram)
🏢 ClanX
📍 Gurugram