02 Aug
|
Firsthive
|
India
Designation - Data Science Engineer
Location: Bengaluru
Experience: 4–6 years
Function: AI & Data Science
Role Description: We are seeking a Data Science Engineer to build and deploy production ML models and AI features for our CDP platform. You will work in a small, high-ownership AI & Data Science team — building customer segmentation models, entity resolution algorithms, predictive analytics, NLP capabilities, and LLM-powered automation that directly impact how enterprise clients understand and engage with their customers. This is a hands-on engineering role — you build models that ship to production, not notebooks that stay in research.
Key Responsibilities:
- Build and deploy customer segmentation and clustering models (K-Means, DBSCAN, hierarchical) at scale
- Develop entity resolution algorithms — fuzzy matching, blocking strategies, probabilistic scoring — to unify customer profiles across disparate data sources
- Build predictive models — churn prediction, conversion propensity, next-best-action recommendations using classification and regression (XGBoost, Random Forest, logistic regression)
- Design and build LLM-powered features — schema mapping automation, natural language querying, AI-driven insight generation using prompt engineering, RAG pipelines, and structured output extraction
- Build NLP capabilities — text embeddings, semantic similarity, entity extraction, text classification using transformers (BERT or similar)
- Write complex SQL for feature engineering — window functions, sessionization, time-series aggregation, customer behavior features from raw event data on Snowflake/BigQuery
- Integrate ML models into the core platform via APIs (FastAPI) for real-time and batch inference
- Own model lifecycle in production — monitoring, drift detection, retraining, versioning
- Work with data engineering teams to ensure clean, structured training data and feature pipelines
Experience and Skills:
- Python ML stack — scikit-learn, Pandas, NumPy, XGBoost. This is 70% of the work.
- LLM / GenAI — prompt engineering, RAG fundamentals, embeddings, vector similarity search, calling LLM APIs (Claude, OpenAI, or similar) with structured outputs. Not fine-tuning — effective use of APIs.
- SQL — complex feature extraction queries on analytical databases. Window functions, sessionization, time-series aggregation, cost-aware query patterns. Not basic SELECT.
- NLP — text embeddings (sentence-transformers or similar), named entity recognition, text classification, semantic search
- Model deployment — FastAPI or Flask, Docker containerization, REST API serving
- Entity resolution / record linkage — fuzzy matching (Levenshtein, Jaro-Winkler), blocking strategies, probabilistic matching across multiple fields
- Model evaluation — precision/recall trade-offs, cross-validation, A/B testing
Positive to Have:
- LangChain, vector databases (Pinecone, FAISS)
- MLflow or experiment tracking
- Kafka consumers for real-time scoring
- Snowflake ML / BigQuery ML
- Time-series forecasting (Prophet, ARIMA)
- Customer analytics or MarTech platform experience
Pay: ₹1,600,000.00 - ₹2,800,000.00 per year
Benefits:
- Commuter assistance
- Health insurance
- Leave encashment
- Provident Fund
Work Location: In person
📌 Data Science Engineer (India)
🏢 Firsthive
📍 India