30 Jul
|
Indpro
|
Bengaluru
About the role
We build and ship machine learning models that make real decisions in production — forecasting, anomaly detection, ranking, classification, and probabilistic systems that customers and revenue depend on. This is a core modeling role. You will spend most of your time framing problems in modeling terms, choosing and training the right models, and proving they work under real-world conditions — noisy data, distribution shift, class imbalance, and the constant threat of leakage.
We are deliberate about one thing: this role is about understanding models, not just wiring them together. We use LLMs and retrieval where they earn their place, and applied GenAI experience is a genuine plus. But if your experience is primarily prompt engineering, RAG pipelines, and vector-database orchestration, this is not the right role — and we'll both save time by being upfront about that now.
If the most interesting thing you did last year was pick a loss function, catch a subtle leak that was inflating your metrics, or prove a simpler model beat a fancier one, we want to talk to you.
What you'll do
- Own problems end to end, from framing to production. Take an ambiguous business problem, decide whether it's a forecasting / classification / clustering / anomaly / ranking / probabilistic problem, and justify that framing. Wrong framing is the most expensive mistake in ML, and we expect you to get it right.
- Build and train models — classical, probabilistic, and deep. Depending on the problem: tree ensembles (XGBoost / LightGBM / CatBoost), time-series models (ARIMA/SARIMA, Prophet, state-space, TCN, and modern sequence models), probabilistic models (HMM, GMM, Bayesian methods), and deep networks (CNNs, RNN/LSTM, transformers) where they genuinely outperform. Fine-tune models properly (LoRA/PEFT, full fine-tunes, or training from scratch) when that's the right tool — not as a euphemism for calling an API.
- Choose evaluation metrics that fit the problem, and defend them. PR-AUC vs ROC-AUC under imbalance, WMAPE vs MAPE for intermittent demand, calibration when probabilities matter. Report the metric that tells the truth, not the one that flatters the model.
- Engineer features and handle messy reality. Class imbalance (weighting, resampling, threshold tuning),
missing data (principled imputation), multicollinearity, and temporal structure. Know when a feature is quietly leaking the target.
- Prevent, detect, and interrogate leakage. Time-based and grouped splits, walk-forward / rolling-origin validation, and a healthy suspicion of any result that looks too good. We consider I caught a leak a badge of honor, not an embarrassment.
- Validate rigorously and report honestly. Cross-validation strategy appropriate to the data, held-out and locked test sets, bootstrap intervals where useful, and clear separation of durable gains from overfitting. Negative results and killed models are expected and valued.
- Ship and operate what you build. Package, version, deploy, and monitor models — batch and real-time. Track drift (data / concept), automate retraining, and build in rollback. You don't have to be a full MLOps specialist, but your models should survive contact with production.
- Use GenAI where it earns its place. Some of our systems combine classical models with retrieval or LLM components. You'll build these when they're the right answer — with the same evaluation rigor you apply to everything else.
- Communicate and collaborate. Explain modeling choices to engineers, product, and non-technical stakeholders. Tie model quality to business outcomes without letting the business framing paper over weak modeling.
What we're looking for (required)
- Genuine core-ML fundamentals. You can explain, from first principles: the bias–variance trade-off, why a model is overfitting and how you'd know, what a given loss function optimizes, how regularization works, and why your chosen validation scheme is sound for this data.
- A track record of models you actually built and evaluated — with the architecture, the data, the metric, and the before/after numbers. Implemented an ML model for various use cases is not a track record.
We want the specifics.
- Fluency in Python and the modern ML stack: scikit-learn, and at least one of PyTorch / TensorFlow used for real training (not just inference). Solid SQL and data-wrangling (pandas / PySpark or equivalent) for large datasets.
- Sound experimental discipline: appropriate cross-validation, leakage prevention, metric selection, and the judgment to prefer a simpler model when it wins.
- The instinct to quantify. You reach for numbers by default, and you know the difference between a self-reported metric and an externally validated one.
- Solid communication and intellectual honesty — you can defend a modeling decision and, just as importantly, admit when a result doesn't hold up.
Nice to have (preferred, not required)
- Deep-learning depth: architectures built or meaningfully modified from scratch, training-dynamics debugging, or published/research work.
- Probabilistic and Bayesian modeling; time-series beyond the standard toolkit (state-space, hierarchical, foundation models for forecasting).
- Genuine fine-tuning experience (LoRA / PEFT / full) with a clear eval story — and the ability to say precisely what it accomplished over prompting.
- Production MLOps maturity: CI/CD for models, drift detection (PSI / Evidently), automated retraining, canary releases, and observability.
- Applied GenAI / RAG experience — welcome as a complement to modeling depth, not a substitute for it.
- Kaggle, open-source ML contributions, or peer-reviewed publications.
- Domain experience in [your domains — e.g., forecasting, fraud/risk, healthcare, retail, manufacturing].
What this role is not
To keep us both efficient, this role is not a fit if your experience is centered on:
- Prompt engineering and LLM orchestration as the primary skill.
- Building RAG pipelines (chunking, embedding, vector search, retrieval) without underlying modeling work.
- Agentic-workflow wiring (LangChain / LangGraph / CrewAI) as the core of your portfolio.
- Integrating third-party model APIs without training or evaluating models yourself.
Location: INDPRO. Pvt. Ltd., Bengaluru (On-site)
Apply: https://indpro.se/careers/ml-engineer / Send an email with your resume [HIDDEN TEXT]
📌 Lead Machine Learning Engineer (Core Modelling) 7+ yrs (Bengaluru)
🏢 Indpro
📍 Bengaluru