13 Aug
|
Quantorus
|
Navi Mumbai
13 Aug
Quantorus
Navi Mumbai
About the Programme
Developing an enterprise AI platform focused on financial compliance and intelligence.
Role Overview
We are looking for a strong Data Engineer to own the data foundation of the platform. Every model, every AI output, and every compliance decision the system makes depends on data arriving reliably, completely, and on time. You will design and build the ingestion pipelines from all source systems into the data platform, own the pipeline monitoring infrastructure, and work closely with internal IT and operations teams to extract data from complex enterprise source
systems.
Key Responsibilities
Data Discovery & Audit
• Conduct a thorough data audit with internal IT and operations teams — map every data
source needed for the platform, assess what already exists on the data platform, and identify gaps.
• Document data sources, schemas, update frequencies, and quality issues for all relevant datasets
• Raise data gaps and quality risks to the Solutions Architect
Pipeline Design & Build
• Design and build ingestion pipelines from all source systems — ERP, government portals, supplier portals, and banking feeds — into the data platform.
• Design pipelines for both batch and real-time ingestion patterns
• Ensure pipelines are idempotent, resumable, and handle source system failures gracefully without data loss or duplication.
Data Quality & Reliability
• Build pipeline monitoring and alerting so data failures are caught and flagged before they corrupt model training or inference.
• Define and implement data quality checks at the point of ingestion — schema validation, completeness checks, and anomaly detection on incoming data volumes.
• Maintain clear data lineage so the team always knows where a data point came from and when it was last updated.
Collaboration & Handoff
• Work closely with internal IT and automation team who hold institutional knowledge of the source systems — this is not a solo exercise.
• Hand off clean, well-documented datasets to the ML Engineers and LLM Engineer for model training and knowledge base building.
• Support the MLOps Engineer in ensuring production pipelines are stable and monitored post-deployment.
Required Qualifications
Education
•B.E. / B.Tech / M.Tech in Computer Science, Information Technology, or a related field.
Experience
• 5+ years of data engineering experience with at least 2 years working on production pipelines at enterprise scale.
• Demonstrated experience building pipelines from complex enterprise source systems — ERP or equivalent.
• Experience building both batch and real-time / streaming ingestion pipelines.
Technical Skills
• Languages: Python and PySpark; SQL proficiency essential.
• Data Platform: Databricks and Delta Lake — must have hands-on production experience.
• Pipeline Orchestration: Apache Airflow, Databricks Workflows, or equivalent.
• Streaming: Kafka, Spark Structured Streaming, or equivalent for real-time ingestion patterns.
• ERP Integration: Experience extracting data from SAP or equivalent large ERP systems strongly preferred.
• API Integration: REST API consumption for government portal or third-party data feeds.
• Data Quality: Experience with data quality frameworks — Excellent Expectations or equivalent.
• Observability: Pipeline monitoring, alerting, and data lineage tooling.
Preferred Qualifications
• Familiarity with SAP data models
• Prior experience with government API ecosystems — GSTN, ICEGATE, or similar.
• Experience building pipelines that feed ML model training workflows.
• Exposure to Unity Catalog or similar data catalogue and governance tools.
• Prior work in fintech, compliance, or tax technology environments.
Skills:- Python, PySpark, SQL, databricks, Delta Lake, Apache Airflow, Apache Kafka and RESTful APIs
📌 Data Engineer (Navi Mumbai)
🏢 Quantorus
📍 Navi Mumbai