02 Sep
|
Proclink
|
Hyderabad
02 Sep
Proclink
Hyderabad
Position Senior Databricks / Azure Data Lake Engineer Experience 6+ Years Location Hyderabad, Telangana Employment Type Full-TimeRole Overview We are looking for an experienced Senior Databricks / Azure Data Lake Engineer with 6+ years of experience in designing, developing, implementing and supporting scalable data engineering solutions on Microsoft Azure. The candidate will be responsible for building robust data pipelines, data lake architectures and transformation frameworks using Azure Databricks, Azure Data Lake Storage (ADLS Gen2), PySpark and SQL. The role requires close collaboration with Data Engineering, Application Engineering, Business Intelligence, Architecture, DevOps and Data Governance teams to deliver secure, scalable and high-quality data solutions.
The ideal candidate should have strong hands-on experience with Databricks, Delta Lake, PySpark, Azure Data Lake and Azure data services, along with good understanding of data architecture and performance optimisation.Key Responsibilities Azure Databricks Design, develop and maintain scalable data engineering solutions using Azure Databricks.
Develop
PySpark notebooks and production-grade data processing jobs. Build reusable frameworks for batch and incremental data processing. Configure and optimise Databricks clusters and compute resources. Implement appropriate cluster policies, job configurations and access controls.
Troubleshoot
Databricks jobs, performance issues and production failures.
Implement Databricks
Workflows/Jobs for scheduling and orchestration.
Azure Data Lake
Design and manage data solutions using Azure Data Lake Storage Gen2 (ADLS). Implement appropriate data lake folder structures and data organisation strategies. Develop ingestion frameworks for structured and semi-structured data. Implement incremental, full-load and CDC-based ingestion patterns.Work with data formats such as: Parquet Delta JSON CSV Avro Implement data partitioning and lifecycle strategies. Ensure data quality, availability and reliability across data lake environments.
Delta Lake
Strong understanding of Delta Lake architecture and capabilities. Implement ACID-compliant data processing using Delta tables.
Work with: MERGE UPDATE DELETE Time Travel Schema Evolution Optimistic Concurrency Implement incremental processing and CDC pipelines.
Optimise
Delta tables using appropriate partitioning and optimisation techniques. Troubleshoot data consistency and performance issues. PySpark &
- Data Engineering Develop scalable data transformation pipelines using PySpark. Write efficient Spark SQL and DataFrame transformations.
Optimise
Spark jobs for large-volume datasets.
Analyse and resolve: Data skew Shuffle issues Partitioning problems Memory issues Long-running jobs Implement appropriate caching, repartitioning and broadcast strategies. Develop reusable PySpark libraries and frameworks. Perform data validation and reconciliation across source and target systems. Data Ingestion &
- Integration Build data pipelines integrating data from multiple sources, including:
REST APIs Relational databases Kafka Files Cloud storage Enterprise applications Implement batch and near-real-time data ingestion.Work with CDC technologies and incremental data processing. Design resilient ingestion frameworks with appropriate error handling and retry mechanisms.
Azure Data Services
Hands-on experience with relevant Azure services such as: Azure Data Lake Storage Gen2 Azure Databricks Azure Data Factory Azure Key Vault Azure Synapse Analytics Azure Event Hubs Azure Monitor Microsoft Entra ID Experience integrating these services into enterprise data platforms is highly desirable.
Azure Data
Factory / Orchestration Develop and maintain Azure Data Factory pipelines. Implement pipeline orchestration between ADF and Databricks. Develop parameterised and reusable pipelines. Implement dependency management and error handling. Configure monitoring, alerts and retry mechanisms. Manage production scheduling and operational support. Data Quality &
- Governance Implement data validation and reconciliation frameworks. Identify and resolve data quality issues. Implement data quality checks at ingestion and transformation stages. Follow enterprise data governance and security standards. Maintain metadata, data lineage and technical documentation. Implement appropriate access controls and data protection mechanisms.
Performance Optimisation
Analyse and optimise Databricks/Spark workloads. Optimise SQL queries and data transformations. Tune Spark configurations and cluster sizing. Identify bottlenecks in data pipelines. Optimise storage and compute costs. Implement appropriate partitioning, caching and file-size management strategies. CI/CD &
- DevOps Integrate Databricks development with CI/CD pipelines.
Experience with Azure DevOps and/or GitHub Actions. Implement source control for notebooks, code and configuration.
Automate deployment across: Development QA UAT Production Implement environment-specific configuration management. Follow enterprise release and change-management processes.
Security
Implement secure access to Azure Data Lake and Databricks. Solid understanding of Azure RBAC and Microsoft Entra ID. Manage secrets using Azure Key Vault. Implement secure service-to-service authentication. Follow enterprise security and compliance requirements.
Experience with Databricks access controls and Unity Catalog is highly desirable.
Production Support
Provide L2/L3 support for production data pipelines. Monitor scheduled jobs and resolve failures within agreed SLAs. Perform root-cause analysis for production incidents. Implement permanent fixes and preventive actions. Participate in on-call/support rotations where required. Maintain operational runbooks and technical documentation.
Required Technical Skills Must
Have 6+ years of experience in Data Engineering. Strong hands-on experience with Azure Databricks. Strong expertise in PySpark. Strong SQL skills.
Hands-on experience with Azure Data Lake Storage Gen2. Strong understanding of Delta Lake.
Experience developing large-scale ETL/ELT pipelines.
Experience with Azure Data Factory.
Experience with batch and incremental data processing.
Experience with data partitioning and performance optimisation. Strong understanding of data engineering principles.
Experience with Git and CI/CD.
Experience supporting production data platforms. Good to Have Databricks certification.
Unity
Catalog experience.
Azure
Synapse experience. Kafka / Event Hubs experience. CDC implementation experience.
Azure
DevOps / GitHub Actions. Terraform / Infrastructure as Code. Python development experience.
Experience with data governance and lineage.
Experience in FinTech / Banking / Financial Services.
Experience handling large-scale transactional datasets.Preferred Technical Stack Cloud : Microsoft Azure Data Platform : Azure Databricks Data Lake : ADLS Gen2 Processing : Apache Spark / PySpark Storage Format : Delta Lake / Parquet Orchestration : Azure Data Factory / Databricks Workflows Database : SQL Server / PostgreSQL / Azure SQL / Synapse Streaming : Kafka / Azure Event Hubs Security : Entra ID / RBAC / Key Vault CI/CD : Azure DevOps / GitHub Actions IaC : Terraform Source Control : GitKey Deliverables The successful candidate will be expected to: Build highly scalable and reliable Azure data pipelines. Develop reusable Databricks/PySpark frameworks. Improve data pipeline performance and reliability.
Implement robust data quality and reconciliation mechanisms. Reduce pipeline processing time and cloud infrastructure costs. Maintain high availability of production data pipelines. Ensure compliance with enterprise security and governance standards.
Improve automation across data engineering processes. Maintain high-quality technical documentation. Collaborate effectively with architects, developers, analysts and business stakeholders.
Experience in Large-Scale Data Platforms Candidates should ideally have experience working with: Large-volume transactional data. Multiple source systems. Complex ETL/ELT pipelines. Batch and near-real-time processing. Production-critical data platforms. Multiple development teams and distributed engineering teams. Enterprise data governance and security standards.
Soft Skills
Strong analytical and problem-solving skills. Excellent communication and collaboration skills. Ability to work with cross-functional engineering teams. Strong ownership of production systems. Ability to troubleshoot complex data and performance problems. Good understanding of software engineering best practices. Strong documentation and knowledge-sharing capabilities. Ability to mentor junior and mid-level data engineers.
Education Bachelors or Masters degree in Computer Science, Information Technology, Engineering, Data Science or a related discipline.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Senior Databricks / Azure Data Lake Engineer (Hyderabad)
🏢 Proclink
📍 Hyderabad