30 Aug
|
NewVison
|
Maharashtra
30 Aug
NewVison
Maharashtra
Role - Data Engineer
Location - Remote(Pan India)
Position Summary
Deloitte is seeking a Data Engineer to support the Omnia Data platform, a fintech data-processing solution that enables users to upload datasets through a web application and receive processed data products and Excel-based reports.
The Data Engineer will design, develop, operate, troubleshoot, and optimize Databricks data pipelines that ingest user-uploaded datasets, apply data transformations and validations, publish data to Unity Catalog, and generate downstream reporting outputs. The role primarily supports Databricks jobs running on serverless compute, while also requiring working knowledge of classic clusters for workload-specific processing, compatibility, troubleshooting, and performance analysis.
The successful candidate will combine strong PySpark and SQL skills with practical experience managing production job runs, diagnosing failures, improving pipeline performance, and working with financial-services data.
Key Responsibilities
Develop scalable data pipelines
- Design and maintain data pipelines using Python, PySpark, Apache Spark, and SQL.
- Build ingestion and transformation processes for datasets uploaded through the Omnia Data web application.
- Implement medallion architecture patterns across Bronze, Silver, and Gold data layers.
- Create reusable, modular, and testable transformation logic for structured and semi-structured data.
- Apply schema validation, data standardization, deduplication, enrichment, and business-rule processing.
- Build curated datasets that support downstream analysis, reporting, and financial-services use cases.
- Ingest processed data into governed tables and objects within Databricks Unity Catalog.
Manage and support Databricks job runs
- Configure, monitor, and maintain Databricks workflows and scheduled job runs.
- Manage task dependencies, parameters, retries, alerts,
execution order, and operational handoffs.
- Investigate failed or incomplete job runs using job logs, Spark UI metrics, query plans, execution details, and data samples.
- Perform root-cause analysis for code defects, data-quality issues, schema changes, dependency failures, and infrastructure-related problems.
- Coordinate reruns, backfills, recovery activities, and production incident resolution.
- Create and maintain operational runbooks, troubleshooting guides, and support documentation.
- Track recurring failures and implement permanent fixes rather than relying on manual reruns.
Apply data governance and quality controls
- Work with Unity Catalog schemas, catalogs, tables, views, volumes, permissions, and metadata.
- Support appropriate access controls for sensitive financial-services and client datasets.
- Implement data-quality checks for completeness, accuracy, uniqueness, validity, consistency, and timeliness.
- Handle schema evolution and unexpected changes in uploaded datasets.
- Maintain explicit metadata, naming conventions, lineage information, and technical documentation.
- Support data traceability from user upload through processing, catalog publication, and final report generation.
- Escalate material data-quality or access issues through the appropriate delivery channels.
Follow engineering and delivery standards
- Write production-quality PySpark and SQL code following established project standards.
- Participate in code reviews, design discussions, backlog refinement,
and technical knowledge-sharing sessions.
- Use version control and project-approved CI/CD practices.
- Develop unit, integration, regression, and data-quality tests where appropriate.
- Collaborate with application developers, QA engineers, product owners, business analysts, and financial-services subject-matter specialists.
- Translate business requirements into technical data-processing rules and acceptance criteria.
- Contribute to continuous improvement of pipeline architecture, operational support, and developer experience.
Required Qualifications
- Bachelor’s degree in Computer Science, Information Systems, Engineering, Mathematics, or a related field, or equivalent professional experience.
- Three or more years of experience in data engineering, data integration, or large-scale data processing.
- Strong programming experience with Python and PySpark.
- Strong working knowledge of SQL, including joins, aggregations, window functions, subqueries, and performance-aware query design.
- Practical experience developing and supporting Databricks notebooks, workflows, and jobs.
- Understanding of Apache Spark concepts such as transformations, actions, lazy evaluation, shuffles, partitions, joins, and execution plans.
- Experience working with Delta Lake and medallion architecture.
- Working knowledge of Unity Catalog and governed data access.
- Experience troubleshooting failed production jobs and analyzing pipeline performance.
- Experience working with data-quality validation, schema management, and reconciliation processes.
- Ability to work effectively in a collaborative delivery environment with distributed technical and business teams.
- Strong written and verbal communication skills, including the ability to document technical issues and explain solutions clearly.
📌 Data Engineer (Maharashtra)
🏢 NewVison
📍 Maharashtra