11 Aug
|
DigitalPaani
|
India
11 Aug
DigitalPaani
India
KEY RESPONSIBILITIES
Build the IoT sensor uncertainty registry, Data Quality Score computations, and cross-sensor + lab validation API/endpoints.
Own GraphDB endpoints for IoT metadata — ingestion, integration, versioning, and materialization of virtual sensors in ClickHouse.
Drive data-quality governance across the platform's sensor and plant datasets.
Build legacy/offline data ingestion pipelines into MongoDB.
Design and maintain ETL pipelines that support IoT-based ML workflows.
SKILLS - MUST HAVE
Robust SQL + Python for production data pipelines
Time-series/IoT data at scale — irregular sampling, gaps, sensor drift, resampling, deduplication.
A columnar/analytical store, ideally ClickHouse (materialized views, projections, query optimisation).
Pipeline/ETL engineering — reliable,
monitored batch + streaming ingestion with idempotency and backfill.
API construction — clean, documented data endpoints.
Data-quality/validation mindset — range/threshold logic, cross-source reconciliation, measurement uncertainty.
Cloud data infra on AWS (EC2/S3 minimum) and a Git/CI-based workflow
SKILLS - NICE TO HAVE
Domain adjacency — water/wastewater, process/industrial, SCADA, PLC/Modbus.
Light MLOps experience — MLflow, SageMaker, model versioning.
MQTT / Node-RED / edge telemetry exposure.
Statistical fluency — uncertainty quantification, error bounds
Graph databases (Neo4j / Cypher) — big plus.
📌 Data Engineer Gurugram (India)
🏢 DigitalPaani
📍 India