Data Engineer (Gurugram)

Data Engineer (Gurugram)

10 Sep
|
Mechademy
|
Gurugram

10 Sep

Mechademy

Gurugram

We are hiring Data Engineer!

Location: Gurugram

Work Mode: Hybrid The Prospect

Our AI platform diagnoses failures on industrial assets worth hundreds of millions of dollars, and every model, copilot, and dashboard sits on top of the data you’d own. We ingest high-frequency sensor and operational data from enterprise clients and turn it into the clean, modeled, reliable datasets that ML and analytics run on.

Your primary job is growing the lakehouse: ingestion from client systems, transformation and modeling, and serving trustworthy data to two consumers, the ML/AI feature pipelines and self-serve analytics for internal teams and clients. Alongside that, you’ll work in the relational database the product runs on, and you’ll need a working understanding of distributed systems to reason about correctness when the pieces between source and lakehouse fail.

About Mechademy

Mechademy builds an enterprise AI platform for real-time monitoring, diagnostics, and predictive maintenance of industrial equipment. Venture-funded and growing fast, 70+ people across New Delhi and Houston, with physics-informed ML and production AI on every deployment. Clients in oil & gas, power generation, and LNG.

What You’ll Own

Lakehouse Pipelines & Ingestion (35%)

- Design and own batch ETL/ELT and CDC pipelines that bring sensor and operational data into the lakehouse, orchestrated in Dagster.
- Build for reliability: idempotent, incremental, backfill-safe pipelines with sane retry and failure handling, that still produce correct output when a worker is killed mid-run or a message is delivered twice.
- Onboard new client data sources: schema and tag mapping, time-series normalization, resampling, gap handling at scale.

Modeling & Serving (25%)

- Model raw data into well-structured, documented tables that downstream ML and analytics can trust.
- Build and maintain the datasets behind ML feature pipelines and the lakehouse layer powering self-serve analytics.
- Write performant Spark/PySpark and SQL; optimize partitioning, storage formats,



and query cost.

Data Quality & Reliability (10%)

- Own data quality: validation, freshness/SLA monitoring, and observability so bad data is caught before it reaches consumers.
- Make the data layer debuggable: lineage, tests, and alerting that tell you what broke and where.
- Reason about failure modes across the whole path (queue, worker, orchestrator, database, object store) and design so that a partial failure leaves the system in a state you can recover from.

Relational & Operational Data (30%)

- Contribute to the schema, indexing, and query performance of the relational database the product runs on.
- Design tables and constraints so that correctness is enforced at the database layer, and diagnose slow queries from their plans.
- Own retention and the boundary between the operational database and the lakehouse: what stays, what moves, and how it gets there.

What Success Looks Like

- First 30 days: Productive in the codebase and orchestration layer. First pipeline change merged.
- First 90 days: Independently shipping and owning pipelines. Onboarded at least one new data source end-to-end.
- First 6 months: Owning a lakehouse data domain, its ingestion, models, and quality, that ML and analytics teams rely on you to drive.

Who You Are Must-Have

- 2–5 years building production data pipelines: real systems with real consumers, not just one-off scripts
- Strong data engineering fundamentals: data modeling, batch vs. streaming, idempotency, incremental processing, partitioning.
- Expert SQL and strong Python: query optimization, window functions, clean production-quality code
- Relational database depth:



you’ve designed schemas for a production PostgreSQL (or equivalent) system and understand normalization and when to break it, indexing strategies, transactions and isolation levels, locking, and how to read a query plan and fix the query
- Distributed systems fundamentals: at-least-once delivery and idempotent consumers, partitioning and its effect on ordering, consistency and durability trade-offs, retries, timeouts, and backpressure. You can explain what happens to in-flight work when a worker or a database node dies
- Hands-on with a distributed processing engine (Spark/PySpark or equivalent) on non-trivial data volumes
- Experience with an orchestrator (Dagster, Airflow, Prefect, or equivalent) and a cloud platform (AWS/Azure)
- Data-quality mindset: you build validation and monitoring into pipelines, not after something breaks

Nice-to-Have

- Time-series or high-frequency sensor data at scale
- TimescaleDB or another time-series database (hypertables, continuous aggregates, compression, retention policies)
- Warehouse/lakehouse modeling (Delta/Iceberg/Snowflake/Redshift or equivalent) and file-format/partition tuning (Parquet)
- CDC / database-replication pipelines
- Message brokers or task queues in production (Kafka, RabbitMQ, or equivalent)
- Building data for ML: feature pipelines, training datasets, serving consistency
- dbt or similar transformation/modeling frameworks
- Docker, Terraform/IaC, CI/CD for data
- IoT, energy, or industrial sector experience. Not required, but it compresses your ramp

Qualifications

- B.Tech / B.E. / B.S. in CS, Engineering, Mathematics, or a related discipline
- 2–5 years of hands-on data engineering experience
- Startup or high-growth environment experience preferred

When you apply, show us what you’ve shipped: a pipeline, a data model, a system real consumers depend on. Be ready to talk through what broke in it and how you found out. That matters more than credentials.

📌 Data Engineer (Gurugram)
🏢 Mechademy
📍 Gurugram

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: data engineer (gurugram) / gurugram