Solution Architect - Technology (Bengaluru)

Solution Architect - Technology (Bengaluru)

30 Aug
|
Omega Healthcare Management Services
|
Bengaluru

30 Aug

Omega Healthcare Management Services

Bengaluru

JOB DESCRIPTION

Principal Data Scientist / SLM Model Development Architect

ROLE SUMMARY

The Principal Data Scientist / SLM Model Development Architect is a senior technical leadership role responsible for defining the architecture, strategy, and enterprise adoption of domain-specific Small Language Models (SLMs), multi-agent systems, and advanced AI solutions. This role provides end-to-end ownership of model design, fine-tuning, optimization, evaluation, deployment, and governance for production AI systems - while mentoring teams, tracking the research frontier, and shaping the organization's applied-AI strategy.

EDUCATIONAL QUALIFICATION

- ME / BE / MCA / PhD
- PhD preferred for advanced research and architectural leadership roles

EXPERIENCE

- 10+ years in Data Science, Machine Learning, and Deep Learning
- Proven experience architecting and deploying production-grade AI/ML and LLM/SLM systems at enterprise scale
- Demonstrated ability to read, evaluate, and rapidly operationalize current research papers into production systems

WORK SCHEDULE

- Day Shift

KEY RESPONSIBILITIES

Architecture & Strategy
- Define and own the enterprise architecture, standards, and roadmap for domain-specific SLM development and multi-agent AI systems.
- Design SLMs optimized for accuracy, latency, cost, and scalability - including on constrained hardware (e.g., single-GPU / small-VRAM / A100/H100 inference and training envelopes).
- Drive architectural decisions across model selection, fine-tuning (QLoRA / LoRA / multi-LoRA / adapter-gated approaches), quantization/compression, and inference optimization.
- Ensure alignment of AI solutions with enterprise security, compliance, and governance requirements.

Advanced Model Development
- Lead the design, development, and deployment of ML, DL, and SLM-based solutions.
- Apply deep mathematical and statistical rigor to model performance - including class-imbalance rebalancing, long-tail coverage, and calibration (avoiding over-emit / over-conservative failure modes).
- Architect efficient model internals and serving using Mixture-of-Experts (MoE), FlashAttention, and sparse / long-context attention techniques.
- Lead distributed and multi-GPU training (DDP / FSDP), including data- and model-parallel strategies, gradient accumulation, and scaling efficiency.
- Own the data engineering pipeline: large-scale, targeted extraction of training and held-out evaluation sets, with leakage-controlled train/eval separation.
- Demonstrate successful delivery of production deployments with measurable business impact.

Multi-Agent Systems & Agentic Workflows
- Architect multi-agent and agentic workflows - orchestration, tool use,



planning, routing, and state management - using frameworks such as LangGraph.
- Design retrieval-augmented systems including RAG and GraphRAG (knowledge-graph-grounded retrieval) for accuracy and traceability.
- Drive harness engineering: build the scaffolding, evaluation loops, guardrails, and orchestration that make agents reliable in production.
- Apply skill optimization - decomposing complex tasks into reusable skills/tools and tuning agent behavior for correctness, cost, and latency.

NLP & Language Modeling
- Architect and guide solutions involving transformers, embeddings, reasoning-trace distillation, and language-model fine-tuning.
- Lead adaptation of base SLMs to domain-specific use cases, including specialization and multi-adapter strategies.
- Establish best practices for prompt/trace design, reward modeling (e.g., GRPO / RLHF with a human reference), evaluation metrics, and continuous performance monitoring.

Inference, Serving & Optimization
- Own inference architecture using modern serving engines such as vLLM and SGLang (continuous batching, paged/prefix KV-cache, structured decoding, multi-adapter serving).
- Drive latency/throughput/cost optimization via quantization, speculative decoding, attention optimization (FlashAttention, sparse attention), compiled/optimized runtimes (e.g., TensorRT / TensorRT-LLM), and hardware-aware tuning.
- Continuously evaluate and adopt recent advancements from the research frontier and translate them into production gains.

Evaluation & Data Quality
- Define the evaluation strategy and reference standards, including human-curated "golden" datasets built under double-blind, multi-annotator adjudication with inter-rater agreement gates (e.g., Cohen's kappa).
- Enforce dataset schema integrity, evidence grounding, cross-component dependencies, and completeness/consistency validation.

Technical & Platform Leadership
- Provide technical authority in Python, scientific computing, and data-processing frameworks.
- Guide teams on PyTorch, Hugging Face Transformers/TRL/PEFT, unsloth/bitsandbytes, and quantized-training tooling; TensorFlow/Keras as applicable.
- Ensure high-quality analytical data access through complex, optimized SQL over large production data stores.




- Oversee development in Linux and GPU-based environments, including constrained-hardware optimization and dependency/environment governance.

Cloud, MLOps & Production Readiness
- Architect and govern ML/AI solutions on AWS or Azure with secure data/artifact handling.
- Define and enforce MLOps standards: CI/CD, model/adapter versioning, checkpointing, monitoring, and lifecycle management.
- Ensure reliability, scalability, and maintainability of AI systems in production agentic orchestration pipelines.

Leadership & Mentorship
- Act as Solution Architect and technical mentor for Data Scientists and ML Engineers.
- Lead design reviews, technical decision forums, and architectural governance boards.
- Foster a research-driven culture - reading, presenting, and operationalizing state-of-the-art papers.
- Collaborate closely with product, engineering, security, and compliance stakeholders.
- Drive innovation while maintaining delivery discipline in Agile environments.

REQUIRED SKILLS & EXPERTISE

- Strong foundation in classical ML: Random Forest, SVM, Regression, Boosting & Bagging
- Deep learning expertise: CNN, RNN, LSTM, GRU, and Transformer architectures
- Advanced NLP and LLM/SLM fine-tuning experience (LoRA/QLoRA/PEFT, quantization, distillation)
- Multi-agent / agentic workflow design with frameworks such as LangGraph
- RAG and GraphRAG (knowledge-graph-grounded retrieval)
- Model efficiency and internals: Mixture-of-Experts (MoE), FlashAttention, sparse / long-context attention
- Distributed / multi-GPU training: DDP / FSDP, data- and model-parallelism
- Inference serving and optimized runtimes: vLLM, SGLang, TensorRT / TensorRT-LLM
- Harness engineering and skill optimization for agentic systems
- Strong track record of reading and operationalizing research papers
- Expert-level Python programming and robust SQL over large datasets
- Cloud-based ML architecture and deployment (AWS/Azure) and MLOps
- Enterprise AI governance, evaluation methodology, and reward/eval-set design

PREFERRED / NICE TO HAVE

- Experience with RLHF/GRPO-style reward modeling and human-in-the-loop golden-dataset programs
- Constrained-hardware / edge-GPU model optimization
- Speculative decoding, structured/guided decoding, and KV-cache optimization
- Contributions to open-source ML/agent tooling or published research
- Computer Vision exposure
- Responsible AI, model risk management, and regulatory compliance
- Experience defining AI strategy at an organizational or platform level

ROLE LEVEL

- Solution Architect
- Individual Contributor with enterprise-level technical leadership responsibilities

📌 Solution Architect - Technology (Bengaluru)
🏢 Omega Healthcare Management Services
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: solution architect - technology (bengaluru) / bengaluru