07 Aug
|
PwC Service Delivery Center
|
Bengaluru
07 Aug
PwC Service Delivery Center
Bengaluru
Responsibilities
- Design, build, and maintain robust ETL/ELT pipelines on Azure using Apache Spark (PySpark and/or Scala) on Databricks/HDInsight
- Orchestrate complex data workflows with Azure Data Factory (pipelines, triggers, integration runtimes) and Databricks Jobs/Workflows
- Develop and manage data lakes on Azure Blob Storage/ADLS Gen2, including naming, partitioning, lifecycle policies, and schema evolution
- Implement curated layers (raw, staged, curated), leveraging columnar formats (Parquet/ORC/Avro) and table formats (Delta Lake); manage metastore/Unity Catalog
- Optimize Spark jobs and cluster configurations for performance and cost (autoscaling, spot VMs, Photon, adaptive query execution, caching, partition tuning)
- Operationalize jobs with monitoring, logging, and alerting via Azure Monitor, Log Analytics, and Databricks metrics; build runbooks and dashboards
- Implement data quality, testing, and observability for pipelines (unit/integration tests, Great Expectations, SLAs, lineage)
- Collaborate with Analytics, Data Science, and Product to deliver modeled, trustworthy datasets for BI, ML, and applications
- Enforce security and governance best practices (AAD RBAC, Managed Identities, ACLs, Key Vault, Private Endpoints, VNet integration, encryption with CMK/SSE)
- Contribute to infrastructure as code and CI/CD (Azure DevOps or GitHub Actions) including Databricks objects deployment
- Participate in on-call rotations, incident response, and postmortems; drive continuous improvement and documentation
Mandatory skill sets
- 6+ years as a Data Engineer (or similar) with a strong focus on Azure data services
- Expert-level experience with Apache Spark (PySpark and/or Scala) and distributed data processing
- Hands-on experience with Databricks and Azure HDInsight for large-scale batch processing
- Proficient in Azure Data Factory for orchestration (pipelines, data flows, triggers)
- Strong Python skills
- Solid SQL (window functions, performance tuning, optimization)
- Practical experience with Azure Blob Storage/ADLS Gen2
- Azure Key Vault
- Azure Monitor/Log Analytics
- Understanding of Hadoop ecosystem fundamentals (HDFS, YARN, Hive/Metastore)
- Strong grasp of data modeling, file formats (Parquet/ORC/Avro), partitioning, and performance best practices
- Experience building production-grade pipelines with testing, monitoring, and alerting
- Version control with Git and cooperative development practices
- Excellent communication and cross-functional collaboration skills
Preferred skill sets
- Stream processing and event-driven architectures (Kafka, Azure Event Hubs, Azure Functions)
- Lakehouse technologies (Delta Lake, Unity Catalog) and query engines (Synapse Serverless, Databricks SQL)
- Governance and lineage tools (Microsoft Purview, OpenLineage)
- Cost optimization on Azure (cluster policies, Photon, spot VMs, storage tiers hot/cool/archive)
- Infrastructure as code (Terraform/Bicep) and CI/CD for data workflows (Azure DevOps/GitHub Actions)
- Containerization and orchestration (Docker, AKS) and packaging for Databricks (dbx, DABs)
- Experience integrating with warehouses (Synapse Dedicated SQL Pools, Snowflake) and BI tools (Power BI)
- Security/compliance exposure (PII handling, least-privilege, network isolation)
Years of experience required 5-7 years
Education qualification: BE/B.Tech/MBA/MCA
Education
- Degrees/Field of Study required: MBA (Master of Business Administration)
- Degrees/Field of Study preferred: Bachelor of Engineering
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 IN_Senior Associate_Data Engineer_Emerging Businesses (Bengaluru)
🏢 PwC Service Delivery Center
📍 Bengaluru