About the role :We are scaling our geospatial data-science capability - turning multi-source data (mobility, points of interest, demographic datasets, transaction signals, satellite imagery) into validated location attributes at fine spatial granularity (grid- and address-level) and powering ML models that are served as real-time APIs. Think: uncovering "the why behind the where."You will own the core data-science work: engineering location features from heterogeneous sources, building geospatial ML models (site selection, sales forecasting, catchment and propensity), cleaning and fusing data, and working with our full-stack engineer to ship models and attributes into production APIs.This is a build role, not a maintenance role. You will help define how our attribute layer, modelling approach, and feature store come together.What you'll do :- Engineer location attributes from heterogeneous sources (mobility/smartphone, POI, demographic, web listed, transactions, satellite) at grid and address-level granularity.- Build, validate, and productionize geospatial ML models: site scoring, demand/sales forecasting, trade-area and catchment analysis, consumer/segment propensity.- Design data-quality pipelines that detect and correct bias, anomalies, and missing data, and keep attributes fresh and accurate.- Establish spatially-aware validation (avoiding spatial leakage) so models generalise across cities and geographies.- Partner with the full-stack engineer to expose models and attributes as real-time, address-level APIs and contribute to a reusable feature store.- Translate ambiguous business questions (site selection, expansion, risk)
into modellable problems and defensible insights for enterprise clients.Primary skills :- Geospatial data science: Robust command of spatial concepts and workflows: coordinate systems/projections, spatial joins, grid/indexing systems (H3, geohash, S2), spatial statistics (spatial autocorrelation / Moran's I), and catchment/trade-area analysis.- Python geospatial stack: Hands-on with GeoPandas, Shapely, Rasterio, GDAL/OGR, and spatial SQL via PostGIS (or BigQuery GIS). Comfortable manipulating vector and raster data at scale.- Machine learning for tabular/spatial problems: Solid grounding in regression and classification, gradient boosted trees (XGBoost/LightGBM), and feature selection, applied to problems like site scoring, demand/sales forecasting, and propensity.- Large-scale feature engineering: Ability to design and generate location attributes from heterogeneous raw sources, and to reason about a feature store - versioning, reuse, freshness - as the backbone of the work.- Data fusion, hygiene, and geocoding: Integrating messy, heterogeneous datasets; imputation, anomaly/bias detection, deduplication and entity resolution; robust geocoding and address/lat-long normalisation.- Performance & spatial-query optimisation: Processing very large point/grid datasets efficiently: spatial indexing (R-tree / GiST), optimised spatial joins, partitioning, query-plan diagnosis,
and geometry simplification - to control runtime and cost.- Big-data and pipeline fluency: Advanced SQL plus distributed processing for large spatial workloads (Spark or Dask), and building reliable, repeatable data pipelines.- Productionizing models: Experience turning models into deployable, real-time APIs in collaboration with engineering - clean, tested, well-documented code (Git) and an understanding of latency, monitoring, and reproducibility.Secondary skills :- Mobility & foot-traffic analytics: Working with smartphone/mobility data for catchment, footfall, and movement patterns.- NLP for unstructured/web-listed data: Extracting structure from text-based sources (listings, reviews, POI descriptions).- Geospatial visualisation: kepler.gl, deck.gl, Plotly, or Streamlit for interactive, map-based storytelling and internal tooling.- Statistics & causal inference: Econometrics, uplift/causal methods, forecasting (time series).- Remote sensing / satellite imagery: Computer vision on imagery (CNNs), Google Earth Engine, land-use classification, building-footprint extraction, NDVI/change detection.- Domain knowledge: Retail/CPG site selection, BFSI credit risk / NPA reduction, e-commerce, or insurance use cases.- Stakeholder communication & B2B product sense: Explaining models and trade-offs to non-technical enterprise clients and shaping the product.Qualifications :- 6+ years of applied data-science experience, with at least one project involving geospatial or location data end-to-end.- Degree in Computer Science, Statistics, Geoinformatics/GIS, Physics, Engineering, or a related quantitative field - or equivalent demonstrable experience. (ref:hirist.tech)
📌 Zoop.One - Senior Data Scientist - Geo Spatial (Pune)
🏢 Zoop
📍 Pune