Data Scientist (Hyderabad)

Data Scientist (Hyderabad)

15 Sep
|
Data Unveil
|
Hyderabad

15 Sep

Data Unveil

Hyderabad

Position Title: Data Scientist — Patient Outcomes & Next Best Action

Experience: 3+ Years

Location: Hyderabad, Telangana

Hire Type: Full Time, On-site

Start Date: Immediate

About the role

We build the analytics and intelligence layer that our life sciences clients and their field teams use to act on patient data — hub enrollments, benefits and prior authorization activity, dispense history, and field engagement, brought together into one view of the patient journey.

We're hiring a Data Scientist to own the predictive models behind that layer. The core question is not "how did the program perform last month?" — it's "what should happen next for this patient, and who should do it?" You'll build the risk, propensity, and recommendation models that tell a Field Reimbursement Manager which cases are about to stall, which patients are drifting toward discontinuation, and which intervention is most likely to work.

This is not a dashboard-building role and it is not a data engineering role. Pipelines and data aggregation sit upstream.

What you'll do

- Model the patient journey end to end — enrollment through benefits verification, prior authorization, first fill, and ongoing therapy — identifying the stages, delays, and drop-off points that predict downstream outcomes.
- Build and own production models for adherence and discontinuation risk, therapy-switch and lapse propensity, time-to-first-fill, and prediction of which access cases will stall in copay or prior authorization.
- Develop Next Best Action recommendations for field reimbursement and field teams: which patients or cases to prioritize, which intervention to take, and when — with the reasoning attached, not just a score.
- Extend the same approach to HCP-level targeting, segmentation, and channel propensity — which provider to engage, on what theme, through which channel.
- Design the explainability surface: surface top contributing factors for every score in language a field user can act on, and make sure the explanation actually reflects the model.
- Build the measurement side honestly — define what a "successful" recommendation is, instrument the feedback loop from field outcomes back into model quality, and evaluate lift against realistic baselines rather than against doing nothing.
- Own the full model lifecycle: feature definition and lineage, training and versioning, deployment, drift and performance monitoring, and scheduled retraining.
- Work directly with product and client-facing teams to turn ambiguous asks into scoped analytical problems with defensible answers.

Technical skills

Statistical foundations

- Survival and time-to-event analysis — Kaplan-Meier, Cox proportional hazards, accelerated failure time, discrete-time hazard models; handling right-censoring and competing risks correctly, which most adherence and discontinuation questions require.
- Longitudinal and panel methods — mixed-effects and hierarchical models, GEE, repeated-measures designs; modeling patients nested within practices, territories, and payers without treating those observations as independent.
- Causal inference — propensity score methods,



inverse probability weighting, difference-in-differences, instrumental variables, and regression discontinuity; understanding confounding, selection bias, and immortal time bias in observational program data.
- Uplift and heterogeneous treatment effect modeling — two-model and transformed-outcome approaches, uplift trees, causal forests, meta-learners (S/T/X-learner); the difference between predicting who will churn and predicting who is persuadable.
- Experiment design — power analysis, sample size determination under small-N conditions, sequential testing, cluster randomization, and interference between units in a field-team context.
- Core inference — hypothesis testing, confidence intervals, multiple comparison correction, bootstrapping, Bayesian estimation and credible intervals, and honest handling of missing-not-at-random data.

Machine learning

- Supervised learning at production quality — gradient boosting (XGBoost, LightGBM, CatBoost), regularized regression, random forests; disciplined hyperparameter tuning, cross-validation schemes that respect temporal ordering, and calibration of predicted probabilities (Platt scaling, isotonic regression) so scores mean what they claim.
- Class imbalance — resampling, class weighting, threshold optimization, and choosing evaluation metrics (PR-AUC, lift at k, expected value framing) appropriate to rare-event prediction.
- Sequence and event-history modeling — sequence classification, embeddings over event streams, recurrent or transformer-based architectures where they earn their complexity over well-engineered tabular features.
- Recommendation and ranking — learning-to-rank, contextual bandits and Thompson sampling for exploration/exploitation in action selection, constrained ranking under field-capacity limits.
- Unsupervised methods — clustering and segmentation (k-means, hierarchical, HDBSCAN, latent class analysis), dimensionality reduction, and anomaly detection over patient and case trajectories.
- Feature engineering on longitudinal data — window aggregations, gap and recency features, trajectory shape features, point-in-time correctness, and rigorous avoidance of target leakage from post-outcome fields.
- Explainability — SHAP, permutation importance, partial dependence and ICE, surrogate models, counterfactual explanations; knowing where each is misleading.

ML engineering & tooling

- Python at production standard — pandas, NumPy, scikit-learn, statsmodels, and at least one of lifelines/scikit-survival, PyTorch, or econML/CausalML; clean, tested, reviewable code rather than notebook sprawl.
- SQL fluency against large relational datasets — window functions, complex joins, CTEs,



query performance awareness; comfort building point-in-time correct training sets from transactional history.
- Cloud ML deployment (AWS preferred) — containerized model serving, batch and near-real-time scoring, model registry and versioning, orchestration (Airflow, Step Functions, or equivalent), and CI/CD for model artifacts.
- Monitoring in production — data drift and concept drift detection, population stability index, feature distribution monitoring, retraining triggers, and rollback strategy when a model degrades.
- Version control and reproducibility — Git, experiment tracking (MLflow, Weights & Biases, or equivalent), and environments where a result from six months ago can be reproduced.

Applied AI

- Familiarity with LLM application patterns — RAG, structured-output prompting, text-to-SQL — and where an LLM is the wrong tool for a scoring problem.
- Generating narrative explanations grounded in model output, and evaluating those narratives for faithfulness rather than fluency.
- Awareness of evaluation practice for generative systems: golden datasets, rubric-based scoring, and regression testing of prompts.

Experience & background

Required

- 3+ years applying machine learning and statistics to production problems, not just analysis notebooks.
- Advanced degree in statistics, biostatistics, epidemiology, computer science, economics, or a related quantitative field — or equivalent demonstrated depth.
- Experience building models on patient- or member-level longitudinal data.
- Ability to explain a model's reasoning and its limits to a non-technical stakeholder without hedging into uselessness.

Strongly preferred

- Healthcare, specialty pharmacy, or life sciences experience — hub services, patient support programs, field reimbursement workflows, advantages investigation and prior authorization, or claims.
- Experience building models that drive human decisions rather than automated ones, and designing for human-in-the-loop review.
- Exposure to patient-level data governance — de-identification, tokenization, HIPAA constraints, and consent and aggregation rules — and comfort designing analyses within those limits.
- Experience contributing to product requirements alongside engineering and design teams.

What success looks like

First 90 days — You understand the patient journey data end to end, know which signals are reliable and which are artifacts of reporting lag, and have a first risk model validated against historical outcomes.

First year — Field teams work their queue in the order your models suggest, trust the explanations enough to repeat them to clients, and there's measurable evidence that acting on the recommendations changes outcomes.

Why this role

Specialty patient data is messy in ways that make it genuinely interesting: fragmented sources, inconsistent reporting cadence, small populations, censored outcomes, and real consequences on every case. You'll have unusual latitude to decide how the modeling problem gets framed, and your work reaches the people making the calls rather than sitting in an internal report.

Industry

- IT Services and IT Consulting

Employment Type

- Full-time

📌 Data Scientist (Hyderabad)
🏢 Data Unveil
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: data scientist (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: data scientist (hyderabad) / hyderabad