24 Sep
|
Finolex Industries (FIL)
|
Pune
24 Sep
Finolex Industries (FIL)
Pune
2.1 Data Platform Architecture &
- Engineering Leadership
•
Technically manage the end-to-end Azure Data Warehouse, Databricks, and Power BI solutions, ensuring high availability, scalability, and performance.
•
Design and own enterprise data models, schemas, and database structures (dimensional, normalized, Data Vault 2.0) to support analytical and operational use cases.
•
Lead the end-to-end implementation and continuous evolution of the Data Lake / Lakehouse architecture using the Medallion model (Bronze, Silver, Gold layers).
•
Define and enforce architectural standards, design patterns, and reference architectures for data ingestion, transformation, storage, and consumption.
•
Evaluate and recommend data storage solutions including relational databases, NoSQL databases, data lakes, lakehouse, and cloud storage services.
•
Identify alternate components within the existing Data Lake Architecture to improve performance and reduce cost through POCs and feasibility studies.
Data Lake Architect
Confidential – Internal Use Only Page 2 of 5
2.2 Data Engineering, Pipelines &
- SAP Integration
•
Design, build, and optimize end-to-end ETL/ELT data pipelines for ingesting, processing, and transforming large volumes of structured, semi-structured, and unstructured data within the Medallion and Lakehouse architecture.
•
Engineer SAP data integration pipelines (ECC / S/4HANA / BW) leveraging tools such as SAP CDS Views, SLT, OData, BAPIs, and third-party replication tools.
•
Implement data validation, quality checks, reconciliation, and observability frameworks (e.g., Great Expectations, Delta Live Tables expectations) to ensure accuracy and consistency.
•
Optimize pipelines and Spark/Databricks workloads (cluster sizing, partitioning, Z-ordering, caching, photon engine) to reduce Azure compute / storage cost and improve latency.
•
Design and maintain CI/CD pipelines using Azure DevOps for automated code migration, testing, and deployment across Dev UAT Prod environments.
•
Manage routine incident handling, root-cause analysis, and service requests within the defined SLAs.
2.3 Analytics, BI &
- AI/ML Enablement
•
Define the semantic layer and enterprise data models in Power BI / Microsoft Fabric to enable self-service analytics for business users.
•
Govern Power BI workspaces, datasets, dataflows, row-level security (RLS), capacity (Fabric F-SKU), and DAX optimization.
•
Partner with Data Scientists and AI engineers to enable advanced analytics, Machine Learning, and Generative AI use cases on the platform (Azure ML, Databricks ML, Azure OpenAI, Mosaic AI).
•
Establish feature engineering pipelines, feature stores, and MLOps practices (MLflow, Azure ML Pipelines) for productionizing models.
•
Drive adoption of Generative AI / RAG solutions on enterprise data, including vector store design, prompt engineering oversight, and responsible AI guardrails.
2.4 Vendor &
- Developer Coordination
•
Act as the technical owner and single point of contact for external vendors, implementation partners, and internal development teams working on the data platform.
•
Review vendor deliverables, technical designs, code, and solution proposals to ensure adherence to architectural standards and best practices.
•
Define and track development milestones, conduct technical reviews, code reviews, and ensure timely delivery of vendor-managed work streams.
•
Drive estimation, planning, and prioritization of data initiatives in collaboration with business and IT leadership.
2.5 Cloud Cost Optimization &
- Budget Management
•
Prepare, present, and track the annual Azure cloud infrastructure budget for data platform components, including compute, storage, networking, and licensing.
•
Continuously monitor Azure consumption using tools such as Azure Cost Management, Azure Advisor, and custom dashboards; identify and act on cost-saving opportunities.
•
Implement FinOps best practices including right-sizing of clusters, reserved/savings plans, spot instances, auto-pause/auto-scaling policies, storage tiering, and lifecycle management.
•
Publish monthly cost reports and variance analysis to leadership with actionable recommendations.
•
Establish chargeback / showback mechanisms for business units consuming data platform services.
2.6 Data Security, Governance &
- Compliance
Data Lake Architect
Confidential – Internal Use Only Page 3 of 5
•
Drive continuous enhancement of data platform security, including identity & access management (Entra ID), network isolation (VNet, Private Endpoints), encryption (at-rest and in-transit), Key Vault, and secure secrets handling.
•
Implement and enforce data governance policies covering data classification, lineage, cataloging, masking, and retention using Microsoft Purview and Databricks Unity Catalog.
•
Ensure compliance with applicable regulatory, audit, and internal information security requirements.
•
Conduct periodic security reviews, vulnerability assessments, and remediation tracking for the data estate.
2.7 Data Integration &
- API Development
•
Build and maintain integrations with internal and external data sources and APIs, including SAP, MES, IoT, CRM, and SaaS applications.
•
Design and implement RESTful APIs, web services, and event-driven interfaces for data access and consumption by downstream applications.
•
Ensure compatibility and interoperability between different systems and platforms.
2.8 Collaboration, Documentation &
- Stakeholder Management
•
Collaborate with business owners, analysts, data scientists, and other stakeholders to understand data requirements and deliver tailored solutions.
•
Prepare Business Requirement Documents (BRDs), High-Level / Low-Level Designs (HLD/LLD), data flow diagrams, and operational runbooks for all current ETL and data platform initiatives.
•
Document technical designs, workflows, and best practices to facilitate knowledge sharing and maintain a current system documentation repository.
•
Provide technical guidance and mentoring to team members and stakeholders as needed.
•
Communicate effectively with business users, the analytics team, data scientists, architects, and other internal and external stakeholders.
3.
Mandatory
Requirements
3.1 Core Technical Skills
•
Proven hands-on experience in data engineering, data architecture, or related roles on the Microsoft Azure stack.
•
Strong proficiency in Python, PySpark, and Spark-SQL; working knowledge of Scala is a plus.
•
Comprehensive expertise in database systems, data modelling methodologies (dimensional, normalized, Data Vault 2.0), and advanced SQL scripting & performance tuning.
•
Hands-on expertise with Azure Data Factory, Azure Databricks, Azure Data Lake Storage (ADLS Gen2), Delta Lake, Azure Synapse / Microsoft Fabric, and Databricks SQL Warehouse.
•
Strong working knowledge of Power BI semantic models, DAX, dataset optimization, RLS, dataflows, and workspace / capacity governance.
•
Sound understanding of SQL and NoSQL databases — preferably SAP HANA, MySQL, MS SQL Server, Azure Cosmos DB.
•
In-depth understanding of Azure DevOps for building CI/CD pipelines, branching strategies, and infrastructure-as-code deployments.
•
In-depth expertise across the Azure cloud platform, including networking, security, monitoring, and cost management services.
3.2 Functional &
- Behavioural Skills
•
Strong problem-solving skills with attention to detail.
Data Lake Architect
Confidential – Internal Use Only Page 4 of 5
•
Excellent communication and stakeholder management skills.
•
Ability to manage multiple priorities, vendors, and project streams concurrently.
•
Ability to adapt to evolving technologies and changing business requirements.
📌 Data Engineer (Pune)
🏢 Finolex Industries (FIL)
📍 Pune