31 Jul
|
Hunting Cherry
|
India
31 Jul
Hunting Cherry
India
ABOUT THE ROLE :
As a Senior Data Engineer on the Data Science team, you will be embedded with data scientists and ML engineers to build the data infrastructure that powers Veho's models, experiments, and operational decisions.
You will spend most of your time in DBT and Databricks designing models, building features, and shipping pipelines that turn raw operational signals into trusted, production-grade datasets the team relies on every day.
Orchestration runs on Airflow, with Prefect in places, and you will own these workflows end-to-end.
RESPONSIBILITIES :
- Design and maintain DBT models that produce trusted datasets, features, and metrics for analysis, experimentation, ML, and reporting.
- Build and operate pipelines in Databricks PySpark jobs and Delta/Iceberg tables that turn raw operational events into analysis-ready data.
- Develop deep familiarity with Veho's operations so the datasets, schemas, and models you ship reflect how the business actually works.
- Orchestrate end-to-end data workflows in Airflow (and Prefect where it fits) with SLAs the DS team can count on for daily models, dashboards, and operational decisions.
- Participate in peer code reviews and raise the bar on data quality, testing, and documentation.
- Partner with data scientists to scope, design, and productionise feature pipelines and model-supporting data.
- Optimise DBT and Spark workloads for cost, performance, and reliability as Veho's data volume grows.
WHAT YOU BRING :
- 8+ years of experience building, testing, and deploying data engineering systems.
- Experience with at least one distributed data system, and the ability to reason about consistency, latency, throughput, and fault tolerance.
- Strong SQL and proficiency with at least one of (py)Spark, DBT, or Airflow in production.
- Experience with Infrastructure-as-Code systems such as Terraform, AWS CDK, or Pulumi.
- Understanding or robust interest in supply chain and the data challenges it creates.
- A self-starter who takes initiative, moves fast, and ships while collaborating on big challenges.
- Excellent written and oral communication in English.
- Comfortable using modern AI coding assistants (Claude Code, Cursor, Copilot, or similar) and experienced with AI-native workflows prompting, agentic tooling, evaluations, retrieval.
- Tech: DBT, Databricks, (py)Spark, Airflow, Prefect, SQL, Iceberg/Delta; familiarity with the broader Veho stack (Kinesis, EMR, Sigma, Pulumi) a plus.
NICE TO HAVE :
- Experience with data lakehouse architectures (Iceberg or Delta).
- Familiarity with Kinesis or other streaming systems in production.
- Exposure to MLOps workflows or supporting ML/AI teams.
- Prior experience working with US-based engineering teams across time zones.
- Demonstrated, measurable success building with LLMs, evaluating model outputs, or integrating AI into data pipelines or internal tooling.
📌 Senior Data Engineer - ETL/PySpark (India)
🏢 Hunting Cherry
📍 India