Clinical AI Data Specialist (HSR Layout)

Clinical AI Data Specialist (HSR Layout)

21 Aug
|
PinnacleU
|
HSR Layout

21 Aug

PinnacleU

HSR Layout

Position: Clinical AI Data Specialist

Location: Bangalore

Working Days: Monday to Friday

Experience Required: 3-7+ years

GROWTH PATH

This is an individual contributor role with strong ownership expectations. High performers may be considered for workstream lead or functional lead responsibilities after approximately 12 months (6–9 months for exceptional candidates), based on demonstrated ownership, delivery, technical judgment, mentoring, cross-functional influence, and ability to reduce dependency on the Director of ML.

ABOUT THE ROLE

We are looking for a Clinical AI Data Specialist to bridge clinical data curation and ML development for oncology-focused AI systems. This person will design ML-oriented labeling and abstraction guidelines for real-world evidence, clinical information extraction, patient-trial matching, registry/QI abstraction, and related clinical workflows. Traditional clinical curation guidelines are often not directly suitable for ML , they may be clinically correct but too ambiguous, inconsistent, non-measurable, or difficult for models to learn. Your role will be to work with clinical experts, clinical data teams, Research Engineers, and ML Evaluation Engineers to create labeling guidelines that are clinically valid, operationally clear, and ML-feasible.

This is not a pure clinical curation role and not a pure ML engineering role. It is a clinical + GenAI + data operations role focused on turning clinical/RWE needs into high-quality datasets and guidelines that ML systems can learn from and be evaluated against.

WHAT YOU WILL DO

- Convert clinical/RWE project requirements into clear, structured labeling guidelines for ML and GenAI systems.
- Define annotation schemas, field definitions, allowed values, evidence-span requirements, inclusion/exclusion rules, normalization rules, and ambiguity-handling logic.
- Manage labeling projects end to end:



understand ML requirements, create guidelines, train annotators/reviewers, run pilot batches, review daily output, resolve ambiguity, and deliver curated datasets.
- Daily review clinical team work for each project: check annotation consistency, missed fields, evidence selection, edge cases, guideline adherence, and recurring disagreement patterns.
- Work with the clinical data team to operationalize annotation projects and improve annotation quality over time.
- Partner with Research Engineers to understand what can be reliably extracted from notes, pathology reports, molecular reports, imaging reports, scanned documents, and other clinical sources.
- Partner with ML Evaluation Engineers to design gold datasets, hidden test sets, adjudication workflows, and label quality checks.
- Use GenAI tools for pre-labeling, guideline drafting, consistency review, error clustering, data inspection, and annotation support where appropriate.
- Use basic Python, SQL, Excel/Sheets, and dashboards to inspect datasets, labels, reviewer output, project status, and quality trends.
- Maintain guideline versioning, change logs, examples, counterexamples, edge-case libraries, and project-level annotation documentation.
- Act as the day-to-day bridge between ML and the clinical data team for new client projects, exploratory projects, and production improvements.

WHAT WE EXPECT

- Background in medicine, clinical research, life sciences, pharmacy, oncology data, clinical informatics, RWE, clinical data abstraction,



or healthcare data operations.
- 3–7+ years of relevant experience in clinical data curation, RWE, clinical research, oncology abstraction, registry abstraction, trial screening, clinical data management, or healthcare data operations.
- Ability to read and interpret complex clinical documents and translate them into accurate annotation rules.
- Robust written communication; ability to write guidelines with examples, counterexamples, edge cases, and clear decision logic.
- Basic programming/data skills: Python basics, pandas/CSV/JSON handling, SQL querying, and Excel/Google Sheets for review, QA, and project tracking.
- GenAI literacy: ability to use LLM tools for drafting, pre-labeling, review, summarization, consistency checks, and data workflows while understanding their limitations.
- Basic ML literacy: labels, training data, validation/test sets, precision/recall, overfitting, leakage, model evaluation, and why labels must be measurable and consistent.
- Strong operational discipline for managing annotation projects, reviewer feedback loops, and versioned guideline updates.
- Ability to work cross-functionally with clinicians, annotators, ML engineers, evaluation engineers, and project stakeholders.

NICE TO HAVE

- Oncology experience, especially with treatment lines, biomarkers, staging, response/progression, pathology, radiology, genomics, adverse events, or trial eligibility.
- Experience designing annotation workflows or working with labeling teams using tools such as Label Studio, Prodigy, internal annotation tools, REDCap, spreadsheets, or review dashboards.
- Experience with RWE datasets, clinical evidence generation, cancer registries, QI abstraction, or EHR-based research.
- Comfort with lightweight scripting, SQL dashboards, Streamlit review tools, or similar data inspection workflows.

📌 Clinical AI Data Specialist (HSR Layout)
🏢 PinnacleU
📍 HSR Layout

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: clinical ai data specialist (hsr layout) / hsr layout

Subscribe to this job alert:

Get the latest job offers by email for: clinical ai data specialist (hsr layout) / hsr layout