08 Sep
|
neurogent.ai
|
Gurugram
08 Sep
neurogent.ai
Gurugram
About the Company neurogent.ai builds agentic AI solutions that help organizations automate workflows and improve customer experience across Banking, Healthcare, Insurance, and Financial Services. We develop intelligent AI agents, chatbots, and voice assistants that provide 24/7 support, streamline operations, and drive measurable efficiency gains. Using modern architectures such as RAG and LangGraph, we help teams reduce costs, improve productivity, and accelerate growth.
Position: Databricks Lead (Data Engineering & MLOps)About the Role
We're hiring a Databricks Lead to run the data platform team for a large US enterprise client — a multi-site operator with 1,000+ locations whose analytics and machine learning programs are being built from the ground up. You'll own the lakehouse and the MLOps backbone that sits on it: governed pipelines from a dozen source systems, feature and prediction tables the data science team can trust, and a promotion path that gets models to production without manual handoffs.
This is a hands-on lead role, not a coordination role. You'll write code, set the standards, review the team's work, and sit directly with the client's data science and IT leadership to make architecture calls and defend them. We're looking for someone with founder instincts — comfortable with ambiguity, decisive with incomplete information, and accountable for outcomes rather than tickets.
Key Responsibilities
- Lead the client's Databricks team — set technical direction, plan and sequence delivery, and own platform outcomes jointly with client data science and IT stakeholders.
- Design and build medallion (Bronze → Silver → Gold) pipelines on Databricks and Delta Lake, integrating operational, CRM, HRIS/payroll, web, and finance sources.
- Build reusable feature, training,
and prediction tables with versioning and multi-year history so models are reproducible and retrainable.
- Stand up sandbox / dev / prod separation with Unity Catalog governance, RBAC, secrets management, and least-privilege access.
- Implement CI/CD for data and ML: automated unit, integration, and end-to-end tests, environment parameterization, and a single approval gate before production promotion.
- Own orchestration — time-based, trigger-based, and manual runs — with DAG-level visibility, retries and recovery, and per-pipeline SLAs.
- Build the MLOps layer: MLflow experiment tracking, model registry and versioning, batch inference pipelines, rollback, and drift and pipeline monitoring with alerting.
- Implement data-quality gates that block model runs on schema changes, null spikes, or out-of-range values, and alert instead of failing silently.
- Make governed data easy to reach from Python and R notebooks; productionize data scientists' code without rewriting it from scratch.
- Deliver outputs to BI (Power BI) and downstream APIs, and prepare the platform for low-latency, real-time scoring as those use cases arrive.
- Mentor engineers, run code reviews, write the runbooks, and keep the team unblocked and shipping.
Required Skills & Experience
- 6+ years in data engineering or data platform roles, including time leading a team or owning a platform end to end.
- Deep hands-on Databricks — Delta Lake, Spark, Workflows, Unity Catalog, MLflow — or equivalent depth on Snowflake,
Microsoft Fabric, or BigQuery with the ability to get productive on Databricks quickly.
- Strong Python and SQL, with production experience building and operating ETL/ELT at scale: incremental loads, schema evolution, backfills, and data modeling.
- Hands-on CI/CD for data and ML workloads: Git workflows, automated testing, pinned package environments, secrets handling, no manual deployment steps.
- Experience with orchestration and data-quality tooling (Databricks Workflows, Airflow, dbt tests, or similar) integrated into the pipeline rather than bolted on.
- Working knowledge of a major cloud (Azure preferred; AWS or GCP acceptable), infrastructure-as-code, and compute cost and performance tuning.
- Ability to work directly with senior client stakeholders — run a working session, write a clear decision document, and hold a technical position under pressure.
- Comfortable using AI tools (Claude Code, Cursor, or similar) to write, review, and accelerate engineering work.
- Excellent communication skills and the ability to operate in a remote, cross-functional workplace with US-based teams.
Nice to Have
- Prior founder or founding-engineer experience — someone who has owned outcomes without a safety net.
- ML/AI depth: feature engineering, model training and deployment, drift monitoring; churn, propensity, or forecasting models in production.
- Streaming and real-time serving experience (Structured Streaming, Kafka or Event Hubs, low-latency inference endpoints).
- Familiarity with R-based data science workflows.
- Power BI or an equivalent semantic/BI layer.
- Experience delivering in a services or consulting model for US clients.
- Databricks certifications (Data Engineer Professional, Machine Learning Engineer) or equivalent.
📌 Databricks Lead (Gurugram)
🏢 neurogent.ai
📍 Gurugram