Role & responsibilities
Design, build, and maintain scalable and robust data pipelines for ingesting, processing, and transforming large volumes of structured and unstructured data.
• Develop and optimize high performance Spark jobs using PySpark and Spark SQL within Azure Databricks.
• Implement data storage solutions using Azure Data Lake Storage (Gen2) following medallion architecture (Bronze, Silver, Gold layers) and best practices for data organization.
• Collaborate with data architects, analysts, and business stakeholders to understand data requirements and translate them into technical solutions.
• Implement data security and compliance measures, including access controls, encryption, and data masking within the Azure ecosystem.
• Perform data modeling to create efficient, scalable schemas for both batch and real time analytics.
• Monitor, troubleshoot, and tune data pipelines and Databricks jobs for performance and cost effectiveness.
• Automate deployment and management of data solutions using CI/CD pipelines (Azure DevOps/GitHub Actions) and Infrastructure as Code (IaC)
tools like Terraform or Bicep.
• Establish and enforce data quality checks and data governance standards across the data platform.
• Mentor junior data engineers and promote best practices in software development and data engineering.Mandatory Skills & Qualifications:
• 5+ years of qualified experience in data engineering, with a proven track record of building enterprise grade data solutions.
• 3+ years of hands on experience with Microsoft Azure data services, specifically:
• Azure Databricks: Expert level proficiency in developing, tuning, and debugging Spark clusters and notebooks.
• Azure Data Lake Storage (Gen2): Deep experience in managing data lakes, including directory structure, security (RBAC & ACLs), and performance optimization.
• Medallion Architecture
• Expert programming skills in Python for data engineering tasks (e.g., Pandas, API interactions, unit testing).
• Expert
📌 Azure Data Engineer (Bengaluru)
🏢 CGI
📍 Bengaluru