Lead Data Engineer (Noida)

Lead Data Engineer (Noida)

28 Sep
|
Algoworks
|
Noida

28 Sep

Algoworks

Noida

Job Summary

Role: Lead Data Engineer Location: India, Remote

Experience: 10+ Years

Role overview

We are seeking a hands-on Senior Data Engineer with strong expertise in Azure Databricks and Azure Data Factory to build, optimize, and maintain scalable enterprise data pipelines. The role will focus on high-performance ETL/ELT development, Delta Lake optimization, data processing across Bronze, Silver, and curated layers, and close collaboration with DWH and reporting teams for downstream consumption.

The ideal candidate will act as a senior technical contributor, ensuring reliability, performance, data quality, and maintainability across the data platform.

Key responsibilities

Pipeline Development

- Build and maintain scalable data pipelines using Azure Databricks and Azure Data Factory.
- Implement ingestion and transformation logic across Bronze and Silver data layers.
- Develop batch and incremental data-processing patterns.
- Design reliable and reusable pipeline components for enterprise workloads.
- Monitor and troubleshoot pipeline execution and data-processing issues.

Curated Layer Delta Lake Development

- Implement hydration, merge, and upsert logic using Delta Lake.
- Build and maintain curated datasets aligned with data quality and business requirements.
- Handle late-arriving data and incremental updates.
- Implement reliable data transformation and reconciliation processes.
- Ensure curated datasets are optimized for downstream consumption.

Performance Storage Optimization

- Optimize Delta Lake tables for performance and cost efficiency.
- Select and tune appropriate storage formats such as Parquet and Delta.
- Apply partitioning, compaction, and file-sizing strategies.
- Tune Spark jobs for large-scale distributed data processing.




- Identify and resolve performance bottlenecks across data pipelines and storage layers.

Downstream DWH Collaboration

- Work closely with DWH and reporting teams to support downstream data consumption.
- Provide optimized datasets for reporting and analytical workloads.
- Support data validation and reconciliation with Gold-layer outputs.
- Collaborate with downstream teams to understand data requirements and optimize delivery.
- Ensure consistency and reliability of data consumed by reporting and analytics platforms.

Engineering Best Practices

- Implement basic CI/CD practices for data pipelines.
- Follow coding standards, documentation, and version-control practices.
- Maintain reusable, scalable, and maintainable pipeline code.
- Support production troubleshooting and performance tuning.
- Participate in Agile delivery processes and technical discussions.

Data Quality Production Support

- Implement data validation and quality checks across ingestion and transformation processes.
- Investigate data discrepancies and pipeline failures.
- Perform root-cause analysis and implement corrective actions.
- Support production deployments and resolve data-processing issues.
- Maintain reliability and consistency across enterprise data pipelines.

Required technical skills and competencies





- Robust hands-on experience in data engineering and enterprise data platforms.
- Strong experience building data pipelines on Azure.
- Advanced proficiency in PySpark.
- Hands-on experience with Azure Databricks.
- Strong experience with Azure Data Factory.
- Deep knowledge of Delta Lake tuning and optimization.
- Strong understanding of storage optimization using Parquet and Delta.
- Strong SQL skills for data transformation, validation, and reconciliation.
- Experience working with large datasets and distributed processing.
- Experience implementing batch and incremental processing patterns.
- Experience with hydration, merge, and upsert logic.
- Experience with Git and basic CI/CD pipelines.
- Familiarity with data quality and validation techniques.
- Experience working in Agile delivery environments.

Must have skills

- Azure Databricks.
- Azure Data Factory.
- PySpark.
- Delta Lake.
- SQL.
- Data Engineering.
- ETL/ELT.
- Data Pipeline Development.
- Bronze/Silver/Curated Data Layers.
- Delta Lake Performance Optimization.
- Spark Performance Tuning.
- Parquet.
- Batch and Incremental Processing.
- Git and Version Control.
- Strong analytical and problem-solving skills.

Good to have skills

- Microsoft Fabric.
- Streaming or near real-time data pipelines.
- Data governance tools.
- Metadata management tools.
- Advanced CI/CD practices.
- Experience with Gold-layer development.
- Experience supporting enterprise reporting and DWH platforms.

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Lead Data Engineer (Noida)
🏢 Algoworks
📍 Noida

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: lead data engineer (noida) / noida