30 Sep
|
neurogent.ai
|
Gurugram
30 Sep
neurogent.ai
Gurugram
About Neurogent
Neurogent.ai builds agentic AI solutions that help organizations automate workflows and improve customer experience across Banking, Healthcare, Insurance, and Financial Services.
We develop intelligent AI agents, chatbots, and voice assistants that provide 24/7 support, streamline operations, reduce costs, improve productivity, and accelerate growth. Our solutions use modern architectures including RAG, LangGraph, and agentic AI to solve complex business and operational problems.
About the role
We are looking for a Lead Data Engineer — Data Platform & MLOps to lead the data platform team for a large US enterprise client — a multi-site organization with 1,000+ locations, where analytics and machine learning capabilities are being built from the ground up.
This is a hands-on technical leadership and business consulting role.
You will own the lakehouse and MLOps foundation supporting the client's analytics and AI initiatives — from governed data pipelines and reusable data products to model deployment and production monitoring. However, we are not looking for someone who simply takes business requirements and builds pipelines.
The ideal candidate should be able to:
Understand the business problem → identify process and operational opportunities → challenge the initial approach when necessary → build a data-backed business case → influence stakeholders → define the right technical solution → lead implementation → measure the business outcome.
You will work directly with senior client stakeholders across Data Science, IT, and business/operations teams, make architecture decisions, defend your recommendations, and take ownership of outcomes.This is a hands-on lead role, not a coordination-only position. You will write code, establish engineering standards, review code, mentor engineers, solve technical problems, and remain deeply involved in platform development.
We value founder-like ownership — someone comfortable with ambiguity, willing to challenge assumptions, decisive with incomplete information, and accountable for business outcomes rather than simply completing tickets.
Key responsibilities
- Data Platform & Lakehouse
- Lead the client's data platform team and establish technical direction, engineering standards, delivery priorities, and platform architecture.
- Design and build scalable Bronze → Silver → Gold medallion pipelines.
- Integrate data from operational systems, CRM, HRIS/payroll, web, finance, and other enterprise sources.
- Build reliable ETL/ELT pipelines supporting incremental loads, schema evolution, backfills, historical data, and data reconciliation.
- Build reusable feature, training, and prediction tables with appropriate versioning and multi-year historical data.
- Optimize compute, storage, pipeline performance, reliability, and platform costs.
- Business Consulting & Business Operations
- Understand the underlying business problem behind a technical request rather than simply implementing the requested solution.
- Analyze existing business processes and identify opportunities for automation, efficiency improvement, cost reduction, revenue improvement, or operational optimization.
- Challenge existing processes, assumptions, and proposed technology solutions when evidence suggests a better approach.
- Evaluate whether a technology investment is justified based on business impact, cost, feasibility, and expected ROI.
- Develop data-backed business cases and clearly communicate recommendations to senior stakeholders.
- Influence business and technical stakeholders to adopt improved processes or alternative solutions when appropriate.
- Translate business objectives into data, analytics, AI, and platform requirements.
- Measure and communicate the business outcomes of implemented solutions.
- Act as a trusted technical and business advisor to the client rather than functioning only as an implementation resource.
- Governance & Security
- Establish sandbox, development, and production environments with appropriate separation.
- Implement catalog-level governance, RBAC, secrets management, and least-privilege access.
- Implement governed access to data from Python, R, BI tools, and downstream applications.
- Establish data lineage, metadata management, and appropriate security controls.
- CI/CD & Engineering Standards
- Implement CI/CD for data and ML workloads using Git-based workflows.
- Establish automated unit, integration, and end-to-end testing.
- Implement environment parameterization and secure secrets handling.
- Establish controlled production promotion with a clear approval process.
- Standardize development practices, code reviews, package management, testing, and deployment.
- Orchestration & Data Quality
- Own pipeline orchestration for scheduled, event-triggered, and manual workflows.
- Establish DAG-level visibility, retries, recovery mechanisms, and pipeline SLAs.
- Implement data-quality gates that prevent unreliable data from reaching downstream systems or triggering ML workflows.
- Detect and alert on schema changes, null spikes, unexpected ranges, duplicate records, and other data-quality issues.
- Ensure failures are visible and actionable rather than silently propagating through the platform.
- MLOps & Productionization
- Build and maintain the MLOps foundation supporting the client's Data Science teams.
- Implement experiment tracking, model registry, model versioning, and controlled model promotion.
- Build batch inference pipelines and support real-time/low-latency inference as required.
- Establish model rollback and recovery mechanisms.
- Implement model and pipeline monitoring, including drift detection and alerting.
- Productionize Data Scientists' models and code without unnecessary rewrites.
- Establish reproducible workflows for model training, validation, deployment, and retraining.
- Analytics & Downstream Integration
- Deliver trusted datasets and outputs to Power BI or equivalent BI platforms.
- Support downstream APIs and analytical applications.
- Design the platform to support future real-time and low-latency scoring use cases.
- Work closely with Data Science, BI, IT, and business teams to ensure data products meet actual business needs.
- Technical Leadership
- Mentor and guide data engineers.
- Conduct code and architecture reviews.
- Establish reusable engineering patterns and standards.
- Write technical documentation, architecture decision records, and runbooks.
- Identify technical risks early and proactively resolve blockers.
- Maintain a high engineering bar while keeping the team shipping.
- Work directly with senior client stakeholders and defend technical decisions when challenged.
Required skills and experience
- 6+ years of experience in data engineering, data platforms, or related engineering roles, including experience leading a team or owning a platform end-to-end.
- Deep hands-on experience with at least one contemporary lakehouse or cloud data platform, such as Databricks, Snowflake, Microsoft Fabric, BigQuery, or equivalent.
- Strong hands-on experience with Python and SQL.
- Production experience building and operating ETL/ELT pipelines at scale, including incremental loads, schema evolution, backfills, data modeling, and data reconciliation.
- Strong experience with Spark or another distributed processing engine, with the judgment to determine when distributed processing is actually required.
- Hands-on experience with CI/CD for data and ML workloads, including Git workflows, automated testing, environment management, secrets handling, and controlled production deployment.
- Experience with orchestration and data-quality tooling such as Airflow, dbt, platform-native workflows, Great Expectations, or similar.
- Working knowledge of a major cloud platform, with Azure preferred; AWS or GCP experience is also acceptable.
- Experience with infrastructure-as-code, cloud security, and compute/performance/cost optimization.
- Experience working directly with senior client stakeholders, including running working sessions, presenting architecture decisions, documenting recommendations, and defending technical positions.
- Business consulting and business operations experience, with the ability to understand the underlying business problem, identify process improvement opportunities, challenge technology requests when appropriate, build a data-backed business case, and influence stakeholders to adopt changes that improve efficiency, cost, revenue, or operational outcomes.
- Ability to translate business objectives into technical architecture and measurable delivery outcomes.
- Comfortable using AI-assisted engineering tools such as Claude Code, Cursor, GitHub Copilot, or similar.
- Excellent written and verbal communication skills.
- Ability to work effectively with remote, cross-functional, and US-based teams.
Nice to have
- Founder, co-founder, founding engineer, or early-stage engineering experience.
- Hands-on Databricks experience, including Delta Lake, Databricks Workflows, Unity Catalog, and MLflow.
- Experience migrating platforms between Databricks, Snowflake, Fabric, BigQuery, or equivalent platforms.
- ML/AI experience including feature engineering, model training, deployment, monitoring, and retraining.
- Experience with churn, propensity, forecasting, recommendation, or other production ML use cases.
- Streaming and real-time experience using Structured Streaming, Kafka, Event Hubs, or equivalent.
- Experience with low-latency model-serving or inference endpoints.
- Familiarity with R-based data science workflows.
- Experience with Power BI or an equivalent semantic/BI layer.
- Experience delivering data, AI, or consulting projects for US-based clients.
- Experience working in a services, consulting, or client-facing engineering environment.
- Relevant certifications such as Databricks, Snowflake, Azure Data Engineer, AWS, or GCP.
📌 Lead Data Engineer (Gurugram)
🏢 neurogent.ai
📍 Gurugram