17 Sep
|
ZecData Technology
|
India
17 Sep
ZecData Technology
India
Job Summary
We are looking for an experienced Data Engineer with 4–9 years of hands-on experience in designing, developing, and maintaining scalable data pipelines and data platforms on Microsoft Azure.
The ideal candidate should have strong expertise in Python, SQL, Azure Data Factory, Azure Data Lake Storage, Databricks, data warehousing, ETL/ELT, and big data technologies. The candidate will be responsible for building reliable data solutions, processing large datasets, optimizing data pipelines, and collaborating with Data Scientists, Data Analysts, BI teams, and software engineers.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines and ETL/ELT workflows on Azure.
- Build data ingestion pipelines from databases, APIs, applications, files, and other data sources.
- Develop data transformation and processing workflows using Azure Data Factory and Databricks.
- Work with large volumes of structured and unstructured data.
- Develop and optimize SQL queries, stored procedures, and data transformations.
- Build and maintain Azure Data Lake and cloud-based data warehouse solutions.
- Implement batch and incremental data processing pipelines.
- Perform data cleansing, validation, transformation, and quality checks.
- Optimize pipelines for performance, scalability, reliability, and cost efficiency.
- Monitor data pipelines and troubleshoot production failures and data-quality issues.
- Implement appropriate logging, alerting, error handling, and recovery mechanisms.
- Collaborate with Data Scientists, Data Analysts, BI developers, software engineers, and business stakeholders.
- Follow best practices for data security, governance, access control, and compliance.
- Participate in architecture and technical design discussions.
- Develop reusable data engineering frameworks and components.
- Maintain technical documentation for data pipelines, workflows, and data architecture.
- Participate in code reviews and follow Agile development methodologies.
Required Technical SkillsProgramming
- Strong hands-on experience with Python.
- Strong proficiency in SQL.
- Good understanding of data structures, algorithms, and software engineering principles.
- Ability to write clean, reusable, maintainable, and production-ready code.
Azure Data Engineering
Strong hands-on experience with multiple Azure services, preferably:
- Azure Data Factory (ADF) – ETL/ELT and pipeline orchestration
- Azure Data Lake Storage Gen2 (ADLS) – Data Lake
- Azure Databricks – Data processing and transformation
- Azure Synapse Analytics – Data warehousing and analytics
- Azure SQL Database – Relational data storage
- Azure Functions – Serverless processing
- Azure Event Hubs – Real-time data ingestion
- Azure Key Vault – Secrets and credential management
- Azure Monitor / Log Analytics – Monitoring and logging
- Microsoft Entra ID (Azure AD) – Identity and access management
Candidates should have practical experience working with Azure data services in production environments.
ETL / ELT & Data Pipelines
- Strong experience designing and developing ETL/ELT pipelines.
- Experience with batch and incremental data processing.
- Experience handling data ingestion from multiple sources.
- Knowledge of pipeline scheduling, dependencies, retries, error handling, and monitoring.
- Experience implementing data validation and quality checks.
- Understanding of CDC and incremental-load strategies is preferred.
Azure Databricks / Apache Spark
- Strong hands-on experience with Azure Databricks.
- Experience with Apache Spark / PySpark.
- Understanding of distributed data processing.
- Experience working with large datasets.
- Knowledge of Spark optimization techniques including:
- Partitioning
- Joins
- Caching
- Data skew
- File optimization
- Experience working with Delta Lake / Delta tables is preferred.
Data Warehousing
- Strong understanding of data warehouse concepts and architecture.
- Experience with Azure Synapse Analytics or similar cloud data warehouses.
- Knowledge of:
- Fact and Dimension Tables
- Star Schema
- Snowflake Schema
- Slowly Changing Dimensions (SCD)
- Data Modeling
- Experience optimizing warehouse queries and data models.
Data Lake & Modern Data Architecture
- Hands-on experience with Azure Data Lake Storage Gen2.
- Understanding of Data Lake and Lakehouse architecture.
- Experience working with file formats such as:
- Parquet
- JSON
- CSV
- Avro
- Understanding of partitioning, schema evolution, and metadata management.
- Experience with Delta Lake / Lakehouse architecture is a strong advantage.
Databases
Experience with one or more:
- Azure SQL
- SQL Server
- PostgreSQL
- MySQL
- Oracle
- MongoDB
Strong understanding of relational databases, query optimization, joins, indexing, and stored procedures is required.
DevOps & CI/CD
- Experience with Git/GitHub/Azure Repos.
- Understanding of CI/CD pipelines using Azure DevOps or similar tools.
- Experience deploying Azure Data Factory pipelines through CI/CD.
- Exposure to Terraform, ARM templates, or Bicep is a plus.
- Basic understanding of Docker and containerization is preferred.
Streaming & Real-Time Data
- Experience with Azure Event Hubs, Kafka, or similar streaming technologies is preferred.
- Understanding of real-time and near-real-time data processing.
- Exposure to Spark Structured Streaming is an advantage.
Valuable to Have
- Experience with Microsoft Fabric.
- Experience with Azure Synapse Analytics.
- Knowledge of Delta Lake and Databricks Lakehouse.
- Experience with Azure Purview / Microsoft Purview for data governance.
- Experience with Azure Monitor and Log Analytics.
- Knowledge of Azure security, IAM, and Key Vault.
- Experience with Airflow or other orchestration tools.
- Experience with Terraform/Bicep/ARM.
- Exposure to Power BI and BI/data analytics workflows.
- Knowledge of data governance, lineage, and data quality frameworks.
Required Experience
- 4–9 years of professional experience in Data Engineering, Big Data Engineering, or a related field.
- Strong hands-on experience with Azure Data Engineering.
- Strong proficiency in Python and SQL.
- Hands-on experience with Azure Data Factory, ADLS Gen2, and Databricks.
- Experience developing and supporting production-grade ETL/ELT pipelines.
- Experience working with PySpark/Apache Spark and large datasets.
- Strong troubleshooting and problem-solving skills.
- Ability to work effectively with cross-functional technical teams.
Educational Qualification
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, Data Science, or a related field.
- Equivalent practical experience may also be considered.
Preferred Tech Stack
Python | SQL | Azure | Azure Data Factory | ADLS Gen2 | Azure Databricks | PySpark | Apache Spark | Azure Synapse | Delta Lake | Azure SQL | Event Hubs | Kafka | Azure DevOps | Git | CI/CD | Microsoft Purview | Key Vault | Azure Monitor | Terraform | Microsoft Fabric
Pay: ₹90,000.00 - ₹110,000.00 per month
Benefits
- Provident Fund
Work Location: Remote
📌 Data Engineer - Azure(Exp: 4-9 yrs) (India)
🏢 ZecData Technology
📍 India