30 Aug
|
Omega Healthcare Management Services
|
Bengaluru
30 Aug
Omega Healthcare Management Services
Bengaluru
JOB DESCRIPTION
Principal Data Scientist / SLM Model Development Architect
ROLE SUMMARY
The Principal Data Scientist / SLM Model Development Architect is a senior technical leadership role responsible for defining the architecture, strategy, and enterprise adoption of domain-specific Small Language Models (SLMs), multi-agent systems, and advanced AI solutions. This role provides end-to-end ownership of model design, fine-tuning, optimization, evaluation, deployment, and governance for production AI systems - while mentoring teams, tracking the research frontier, and shaping the organization's applied-AI strategy.
EDUCATIONAL QUALIFICATION
- ME / BE / MCA / PhD
- PhD preferred for advanced research and architectural leadership roles
EXPERIENCE
- 10+ years in Data Science, Machine Learning, and Deep Learning
- Proven experience architecting and deploying production-grade AI/ML and LLM/SLM systems at enterprise scale
- Demonstrated ability to read, evaluate, and rapidly operationalize current research papers into production systems
WORK SCHEDULE
- Day Shift
KEY RESPONSIBILITIES
Architecture & Strategy
- Define and own the enterprise architecture, standards, and roadmap for domain-specific SLM development and multi-agent AI systems.
- Design SLMs optimized for accuracy, latency, cost, and scalability - including on constrained hardware (e.g., single-GPU / small-VRAM / A100/H100 inference and training envelopes).
- Drive architectural decisions across model selection, fine-tuning (QLoRA / LoRA / multi-LoRA / adapter-gated approaches), quantization/compression, and inference optimization.
- Ensure alignment of AI solutions with enterprise security, compliance, and governance requirements.
Advanced Model Development
- Lead the design, development, and deployment of ML, DL, and SLM-based solutions.
- Apply deep mathematical and statistical rigor to model performance - including class-imbalance rebalancing, long-tail coverage, and calibration (avoiding over-emit / over-conservative failure modes).
- Architect efficient model internals and serving using Mixture-of-Experts (MoE), FlashAttention, and sparse / long-context attention techniques.
- Lead distributed and multi-GPU training (DDP / FSDP), including data- and model-parallel strategies, gradient accumulation, and scaling efficiency.
- Own the data engineering pipeline: large-scale, targeted extraction of training and held-out evaluation sets, with leakage-controlled train/eval separation.
- Demonstrate successful delivery of production deployments with measurable business impact.
Multi-Agent Systems & Agentic Workflows
- Architect multi-agent and agentic workflows - orchestration, tool use,
planning, routing, and state management - using frameworks such as LangGraph.
- Design retrieval-augmented systems including RAG and GraphRAG (knowledge-graph-grounded retrieval) for accuracy and traceability.
- Drive harness engineering: build the scaffolding, evaluation loops, guardrails, and orchestration that make agents reliable in production.
- Apply skill optimization - decomposing complex tasks into reusable skills/tools and tuning agent behavior for correctness, cost, and latency.
NLP & Language Modeling
- Architect and guide solutions involving transformers, embeddings, reasoning-trace distillation, and language-model fine-tuning.
- Lead adaptation of base SLMs to domain-specific use cases, including specialization and multi-adapter strategies.
- Establish best practices for prompt/trace design, reward modeling (e.g., GRPO / RLHF with a human reference), evaluation metrics, and continuous performance monitoring.
Inference, Serving & Optimization
- Own inference architecture using modern serving engines such as vLLM and SGLang (continuous batching, paged/prefix KV-cache, structured decoding, multi-adapter serving).
- Drive latency/throughput/cost optimization via quantization, speculative decoding, attention optimization (FlashAttention, sparse attention), compiled/optimized runtimes (e.g., TensorRT / TensorRT-LLM), and hardware-aware tuning.
- Continuously evaluate and adopt recent advancements from the research frontier and translate them into production gains.
Evaluation & Data Quality
- Define the evaluation strategy and reference standards, including human-curated "golden" datasets built under double-blind, multi-annotator adjudication with inter-rater agreement gates (e.g., Cohen's kappa).
- Enforce dataset schema integrity, evidence grounding, cross-component dependencies, and completeness/consistency validation.
Technical & Platform Leadership
- Provide technical authority in Python, scientific computing, and data-processing frameworks.
- Guide teams on PyTorch, Hugging Face Transformers/TRL/PEFT, unsloth/bitsandbytes, and quantized-training tooling; TensorFlow/Keras as applicable.
- Ensure high-quality analytical data access through complex, optimized SQL over large production data stores.
- Oversee development in Linux and GPU-based environments, including constrained-hardware optimization and dependency/environment governance.
Cloud, MLOps & Production Readiness
- Architect and govern ML/AI solutions on AWS or Azure with secure data/artifact handling.
- Define and enforce MLOps standards: CI/CD, model/adapter versioning, checkpointing, monitoring, and lifecycle management.
- Ensure reliability, scalability, and maintainability of AI systems in production agentic orchestration pipelines.
Leadership & Mentorship
- Act as Solution Architect and technical mentor for Data Scientists and ML Engineers.
- Lead design reviews, technical decision forums, and architectural governance boards.
- Foster a research-driven culture - reading, presenting, and operationalizing state-of-the-art papers.
- Collaborate closely with product, engineering, security, and compliance stakeholders.
- Drive innovation while maintaining delivery discipline in Agile environments.
REQUIRED SKILLS & EXPERTISE
- Strong foundation in classical ML: Random Forest, SVM, Regression, Boosting & Bagging
- Deep learning expertise: CNN, RNN, LSTM, GRU, and Transformer architectures
- Advanced NLP and LLM/SLM fine-tuning experience (LoRA/QLoRA/PEFT, quantization, distillation)
- Multi-agent / agentic workflow design with frameworks such as LangGraph
- RAG and GraphRAG (knowledge-graph-grounded retrieval)
- Model efficiency and internals: Mixture-of-Experts (MoE), FlashAttention, sparse / long-context attention
- Distributed / multi-GPU training: DDP / FSDP, data- and model-parallelism
- Inference serving and optimized runtimes: vLLM, SGLang, TensorRT / TensorRT-LLM
- Harness engineering and skill optimization for agentic systems
- Strong track record of reading and operationalizing research papers
- Expert-level Python programming and robust SQL over large datasets
- Cloud-based ML architecture and deployment (AWS/Azure) and MLOps
- Enterprise AI governance, evaluation methodology, and reward/eval-set design
PREFERRED / NICE TO HAVE
- Experience with RLHF/GRPO-style reward modeling and human-in-the-loop golden-dataset programs
- Constrained-hardware / edge-GPU model optimization
- Speculative decoding, structured/guided decoding, and KV-cache optimization
- Contributions to open-source ML/agent tooling or published research
- Computer Vision exposure
- Responsible AI, model risk management, and regulatory compliance
- Experience defining AI strategy at an organizational or platform level
ROLE LEVEL
- Solution Architect
- Individual Contributor with enterprise-level technical leadership responsibilities
📌 Solution Architect - Technology (Bengaluru)
🏢 Omega Healthcare Management Services
📍 Bengaluru