24 Sep
|
protiviti india
|
Coimbatore
24 Sep
protiviti india
Coimbatore
JD: Data Engineer:
Responsibilities
Data Platform Design & Architecture
- Work with Engineering Lead and maintain the firm's enterprise data warehouse on Microsoft Fabric covering Lakehouse, Warehouse, and Semantic layers end-to-end.
- Maintain medallion-pattern data pipelines (Bronze Silver Gold) that consolidate data from practice verticals, internal operations, and client-facing systems.
- Define and enforce logical and physical schema design standards, ensuring data structures are optimized for both analytical queries and downstream AI consumption.
- Maintain architecture decision records, solution blueprints, and platform documentation to enable scalable, auditable infrastructure.
Pipeline Engineering & ETL/ELT Development
- Design, build, and operationalize robust ETL/ELT pipelines using PySpark, Fabric Notebooks, and Fabric Data Factory — handling ingestion, transformation, validation, and load.
- Integrate diverse data sources — CRM, ERP, project management tools, financial systems, and cloud-based SaaS applications — into a unified data platform.
- Implement incremental and batch ingestion patterns; design fault-tolerant pipelines with monitoring, alerting, and self-healing retry logic.
- Build and maintain reusable pipeline templates and modular transformation logic to accelerate future development.
Performance, Reliability & Platform Operations
- Own performance tuning, query optimization, indexing strategies, and capacity planning across the Fabric Lakehouse and Warehouse environments.
- Manage upgrade cycles, schema evolution, and version-controlled deployments using CI/CD best practices integrated with Azure DevOps.
- Implement SLA monitoring and data freshness alerting to ensure pipelines meet business-critical availability requirements.
- Conduct regular platform health reviews and proactively identify and resolve bottlenecks before they impact downstream consumers.
Analytics Enablement & Semantic Layer
- Build and maintain Power BI semantic models (datasets) on top of the Gold layer — ensuring business-friendly naming, consistent measures, and row-level security.
- Enable self-service analytics for consultants and operations teams by creating well-documented, reusable datasets that reduce dependency on the data team.
- Collaborate with analysts and practice leads to translate KPI definitions into reliable, calculated measures within the semantic layer.
- Develop and maintain interactive Power BI dashboards for practice performance, utilisation, revenue, and operational health.
Technology repository
- Document and maintain the firm's data engineering playbooks, onboarding guides, and standards so the platform can scale as the team grows.
Experience, Qualifications & Skills Must-Have
- 4-5 years of hands-on experience in data engineering, data platform development, or related roles.
- Deep, practical expertise with Microsoft Fabric — Lakehouse, Warehouse, Data Factory, Fabric Notebooks, and OneLake.
- Proficiency in PySpark and SQL for large-scale data transformation, aggregation,
and pipeline development.
- Proven experience designing and implementing Medallion (Bronze/Silver/Gold) architectures.
- Strong Python skills for scripting, automation, and data pipeline orchestration.
- Experience with Azure cloud services — Azure Data Lake Storage Gen2, Azure Synapse, Azure Data Factory, Azure DevOps.
- Solid understanding of data warehousing principles — schema design, ETL/ELT best practices, data quality, and lineage.
- Experience building Power BI semantic models and BI dashboards from data engineering outputs.
Good-to-Have
- Familiarity with Snowflake, Databricks, or other cloud-native data platforms.
- Exposure to data mesh architecture principles, data contracts, or federated governance models.
- Experience with streaming / real-time ingestion using Fabric Eventstream, Kafka, or Azure Event Hubs.
- Background in skilled services, consulting, or internal analytics functions.
Key Competencies MS Fabric (Lakehouse / Warehouse)
Azure Data / Synapse
PySpark / Spark
Data Factory / Pipelines
Power BI & Semantic Models
Medallion Architecture
ETL / ELT Design
Python & SQL
Data Governance & Lineage
Schema Design
Performance Tuning
KPI Frameworks
Educational Qualifications
- Bachelor's or Master's degree in Computer Science, Data Engineering, Information Systems, Statistics, or a related quantitative field.
- Relevant certifications valued: Microsoft Certified: Fabric Analytics Engineer Associate, Azure Data Engineer Associate (DP-203), or equivalent.
- A demonstrable portfolio of data platform builds, pipeline architectures, or open-source contributions is a compelling differentiator.
📌 Data Engineer (Coimbatore)
🏢 protiviti india
📍 Coimbatore