Senior Data Engineer (Pune)

Senior Data Engineer (Pune)

22 Sep
|
Bajaj Finance
|
Pune

22 Sep

Bajaj Finance

Pune

Job Summary

The MLOps Engineer will own the end-to-end operationalisation of machine learning, large language model (LLM), and agentic AI workloads on the Bajaj Finance Enterprise Data Platform - a 5PB+ medallion lakehouse built on Azure Databricks and Unity Catalog. This role sits at the intersection of data engineering, model lifecycle management, and AI governance, ensuring that every model - from classical ML to RAG pipelines and autonomous agents - is reproducible, explainable, observable, and production-safe. The incumbent will architect and implement the MLOps and LLMOps platform on Databricks, leveraging Agentbricks (Databricks Agent Framework), Databricks Apps, MLflow, Feature Store, Model Serving, and Mosaic AI - embedding rigorous CI/CD, drift monitoring, cost governance, and responsible-AI guardrails across the full lifecycle. This is a high-impact, high-visibility role critical to delivering Bajaj Finance's AI-first data strategy at scale across 120M+ customer interactions.

Duties and Responsibilities

- Databricks Apps & Self-Serve AI: Develop and deploy internal AI-powered applications using Databricks Apps - enabling business users to interact with ML models, RAG systems, and analytics agents through governed, self-serve interfaces.
- Databricks Apps & Self-Serve AI: Integrate Databricks Apps with Unity Catalog row/column-level security, ensuring data access controls are enforced transparently without requiring users to understand the underlying platform.
- Databricks Apps & Self-Serve AI: Build reusable application templates and deployment blueprints for common use cases (credit decisioning dashboards, collections intelligence tools, KYC automation) to accelerate delivery across business units.
- Monitoring, Observability & Governance: Implement comprehensive model monitoring: data drift (population stability index, KS-statistic), concept drift, prediction drift, and feature distribution shifts - with automated alerts and retraining triggers via Databricks Workflows.
- Monitoring, Observability & Governance: Build model performance dashboards in Databricks SQL / Power BI tracking accuracy, F1, AUC, RMSE, and business KPIs (approval rate, delinquency lift) across all production models.
- Monitoring, Observability & Governance: Enforce Unity Catalog-based data and model lineage - every model must have traceable lineage from raw source data through features to predictions, satisfying RBI Model Risk Management guidelines and BCBS 239.
- Monitoring, Observability & Governance: Conduct regular model validation and bias audits; document model cards and maintain model risk registers in collaboration with Risk and Compliance teams.
- Monitoring, Observability & Governance: Implement cost governance: cluster auto-scaling policies, spot-instance strategies, DBU budget alerts, and compute right-sizing recommendations to optimise the platform spend within approved budgets.
- Collaboration & Engineering Excellence: Partner with Data Scientists, AI Engineers, Data Engineers, and Business stakeholders to productionise models rapidly without sacrificing quality or compliance.
- Collaboration & Engineering Excellence: Establish and evangelise MLOps best practices,



coding standards, and platform conventions through documentation, internal training sessions, and code reviews.
- Collaboration & Engineering Excellence: Contribute to the EDIL technical roadmap - evaluating emerging Databricks capabilities (Delta Live Tables, Lakeflow, Genie Spaces, AI/BI Dashboards) and proposing adoption plans with explicit ROI justification.
- A. MLOps Platform Engineering: Design, build, and maintain the end-to-end MLOps platform on Azure Databricks - covering experiment tracking (MLflow), model registry, Feature Store, batch and real-time model serving, and automated retraining pipelines.
- A. MLOps Platform Engineering: Implement CI/CD pipelines for ML code (Databricks Asset Bundles / DABs, Azure DevOps, GitHub Actions) ensuring reproducible model builds, automated testing, and zero-downtime deployments.
- A. MLOps Platform Engineering: Govern the full model lifecycle: versioning, lineage tracking via Unity Catalog, promotion workflows (Dev Staging Production), and model archival with audit trails.
- A. MLOps Platform Engineering: Establish and maintain Feature Store - curated, reusable feature sets across credit risk, fraud, customer propensity, and collections models - ensuring data freshness, SLA adherence, and lineage traceability.
- A. MLOps Platform Engineering: Operationalise Databricks Model Serving (serverless + provisioned endpoints) and Mosaic AI for scalable, low-latency inference across batch and online serving patterns.
- B. LLMOps - Large Language Model Lifecycle: Design and implement LLMOps pipelines for RAG-based applications on Databricks: document ingestion chunking embedding generation vector indexing (Mosaic AI Vector Search / Unity Catalog Volumes) retrieval LLM serving.
- B. LLMOps - Large Language Model Lifecycle: Implement prompt versioning, prompt evaluation frameworks (MLflow LLM Evaluate, Mosaic AI Evaluation), and automated hallucination / faithfulness / relevance scoring using LLM-as-a-Judge patterns.
- B. LLMOps - Large Language Model Lifecycle: Manage LLM fine-tuning workflows: curate supervised fine-tuning datasets, run PEFT/LoRA jobs on Databricks GPU clusters, register and serve fine-tuned models via MLflow Model Registry.
- B. LLMOps - Large Language Model Lifecycle: Build token-cost monitoring, latency tracking, and model quality dashboards; implement automated rollback triggers when LLM quality KPIs degrade beyond defined thresholds.
- B. LLMOps - Large Language Model Lifecycle: Enforce LLM governance: input/output guardrails, PII redaction, jailbreak detection, and compliance logging aligned with RBI and DPDP Act requirements.
- C. Agentbricks & Agentic AI Operationalisation: Deploy and operationalise autonomous AI agents using Databricks Agentbricks (Agent Framework)



- including tool-calling agents, multi-agent orchestration, and human-in-the-loop review gates.
- C. Agentbricks & Agentic AI Operationalisation: Implement agent observability: trace logging (MLflow Traces), latency profiling, tool-call auditing, and failure mode analysis for production agents such as FinOps Sentinel and Governed Analytics.
- C. Agentbricks & Agentic AI Operationalisation: Build agent evaluation harnesses - synthetic scenario libraries, adversarial test suites, and regression benchmarks - to validate agent behaviour before and after model updates.
- C. Agentbricks & Agentic AI Operationalisation: Manage agent state and memory persistence using Databricks-native storage (Delta Lake, Unity Catalog) and integrate with external databases (CosmosDB, Neo4j) as required by agent workflows.
- C. Agentbricks & Agentic AI Operationalisation: Collaborate with AI Engineers on agent architecture decisions and ensure all agentic workloads meet latency SLAs, cost budgets, and safety standards.
- Key Decisions / Dimensions: Selection of MLOps tooling and pipeline patterns within the approved Databricks platform stack.
- Key Decisions / Dimensions: Model promotion from Staging to Production for models below defined risk thresholds after successful evaluation.
- Key Decisions / Dimensions: Compute cluster configurations, auto-scaling policies, and spot-instance strategies for ML workloads.
- Key Decisions / Dimensions: Drift alert thresholds and automated retraining triggers for registered models.
- Key Decisions / Dimensions: Agent trace sampling rates, logging retention policies, and observability dashboard design.
- Major Challenges: Balancing velocity and rigour: delivering fast model deployments across 50+ source systems and 120M+ customer records while maintaining strict audit trails demanded by RBI Model Risk Management frameworks.
- Major Challenges: LLM non-determinism in production: managing hallucination risk, prompt sensitivity, and output variability in customer-facing AI applications where errors have direct financial and regulatory consequences.
- Major Challenges: Agentic AI safety: ensuring autonomous agents operating on live financial data remain within sanctioned boundaries - particularly for high-stakes decisions such as credit line adjustments, fraud flags, and collections prioritisation.
- Major Challenges: Scale and latency: serving real-time inference (sub-100ms) while maintaining model quality and managing compute costs within budget.
- Major Challenges: Cross-functional alignment: coordinating model deployment gates across Data Science, Risk, Compliance, IT Security, and Business teams - each with different timelines, priorities, and risk appetites.
- Major Challenges: Keeping pace with the Databricks roadmap: the platform evolves rapidly (Agentbricks, Mosaic AI, Lakeflow); the incumbent must continuously evaluate and integrate new capabilities without destabilising production workloads.
- Required Qualifications and Experience: a. B.Tech / B.E. / M.Tech / M.S. in Computer Science, Information Technology, Data Science, Electrical Engineering, or a related quantitative discipline.

📌 Senior Data Engineer (Pune)
🏢 Bajaj Finance
📍 Pune

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior data engineer (pune) / pune