30 Sep
|
Ekshvaku Tech Innovations
|
Hyderabad
30 Sep
Ekshvaku Tech Innovations
Hyderabad
Senior AI/ML Engineer — Clinical AI Quality & LLMOps
Role Purpose
We are looking for a Senior AI/ML Engineer — Clinical AI Quality & LLMOps to establish dedicated ownership of the quality, reliability, evaluation, optimization, and operational maturity of DoctusMind's clinical AI capabilities.
- This role will sit at the intersection of AI/ML Engineering, Clinical AI Quality, LLMOps, AI Evaluation, Prompt Engineering, and Production AI Operations.
- The person will work closely with Product, Engineering, QA, Clinical, and other cross-functional teams to ensure that DoctusMind's AI systems are clinically reliable, measurable, observable, cost-productive, and production-ready.
- The role is not limited to prompt engineering. The expectation is that this person will build the systems, processes, evaluation frameworks, and engineering practices required to operate clinical AI reliably at scale.
- The initial focus will include AI stabilization, clinical AI quality, Chronic Care, and longitudinal intelligence, while establishing the foundation for future specialty, multimodal, agentic, retrieval, memory, and other clinical AI capabilities.
What Success Looks Like The person will be accountable for five executive-level outcomes during the first 3–6 months.
1. AI Quality & Stabilization
2. Production AI Release Governance
3. AI Observability & Operational Performance
4. Clinical AI Capability Production Readiness
Core Responsibilities
- Prompt Engineering & Management
- Design, optimize, and maintain production-grade prompts.
- Establish prompt versioning and lifecycle management.
- Build prompt regression suites.
- Analyze prompt-related failures.
- Optimize prompts for quality, consistency, latency, and cost.
- Maintain traceability between prompt changes and evaluation outcomes.
2. AI Evaluation & Clinical Quality
- Build AI evaluation frameworks for accuracy, relevance, consistency, safety, and reliability.
- Develop golden datasets and clinical test cases.
- Create automated evaluation workflows.
- Establish severity classification for AI failures.
- Perform systematic failure analysis and root-cause analysis.
- Monitor clinical AI quality continuously.
- Partner with Clinical and QA teams to validate high-risk workflows.
3. AI Regression & Release Management
- Establish AI regression testing as part of the release lifecycle.
- Define evaluation gates for prompts, models, retrieval, and configurations.
- Ensure AI changes cannot reach production without appropriate validation.
- Maintain AI version traceability.
- Establish rollback and post-release validation processes.
4. LLMOps & AI Observability
- Build and maintain AI-specific monitoring.
- Track quality, latency, failures, token consumption, and cost.
- Develop dashboards and operational reporting.
- Establish alerts for significant deviations.
- Support AI production incident investigation.
- Identify AI quality and performance drift.
5. AI Experimentation & A/B Testing
- Maintain environments for prompt and model experimentation.
- Conduct controlled A/B tests.
- Compare prompts, models, retrieval strategies, and AI workflows.
- Define experiment methodology and success criteria.
- Document findings and recommendations before production adoption.
6. Model Benchmarking & Selection
- Benchmark existing and emerging models against DoctusMind use cases.
- Evaluate models based on:
- Clinical quality
- Reliability
- Safety
- Latency
- Cost
- Scalability
- Recommend model selection and routing strategies.
- Evaluate fallback strategies for production resilience.
7. Cost & Token Optimization
- Analyze LLM/API consumption.
- Optimize prompt and context length.
- Evaluate caching opportunities.
- Optimize model selection and routing.
- Identify unnecessary or inefficient AI calls.
- Establish ongoing AI cost-performance monitoring.
8. Context Engineering, Retrieval & Memory
- Optimize how patient information, conversations, care plans, and clinical context are supplied to AI systems.
- Improve retrieval of relevant patient history and clinical context.
- Support RAG, embeddings, vector databases, and retrieval pipelines.
- Contribute to cross-session memory and longitudinal intelligence.
- Evaluate context quality and relevance as part of AI evaluation.
9. Clinical AI Capability Productionization
- Evaluate current AI capabilities before production.
- Support Chronic Care and longitudinal intelligence initiatives.
- Contribute to specialty-specific AI capabilities.
- Support multimodal clinical AI capabilities.
- Evaluate AI agents, tools,
reasoning capabilities, and emerging technologies.
- Establish production-readiness criteria for new capabilities.
10. Continuous AI Improvement
- Establish feedback loops using clinician, patient, QA, and system signals.
- Analyze production feedback and recurring failure patterns.
- Convert production learnings into evaluation datasets and regression tests.
- Continuously improve AI quality and operational performance.
Candidate Profile
Required Experience
5+ years of software engineering / ML engineering experience, including:
- 2–3+ years building and operating production LLM/Generative AI systems
- Hands-on experience with AI evaluation and regression testing
- Production LLMOps experience
- AI observability and monitoring
- Model benchmarking and experimentation
- Prompt engineering
- Production AI troubleshooting
- API-based AI/LLM applications
- Python
- AI performance and cost optimization
Healthcare Experience — Required Candidates should have demonstrated experience in one or more of:
- Healthcare AI
- Clinical AI
- Healthcare technology
- Clinical NLP
- Healthcare data platforms
- Clinical decision-support systems
- Healthcare-focused LLM applications
Healthcare/clinical AI experience should be treated as a core screening criterion, not simply a preferred qualification.
Technical Skills
Strong hands-on experience with several of the following:
LLM / GenAI
- OpenAI / Gemini / Claude or equivalent LLM platforms
- Prompt engineering
- Structured generation
- Function/tool calling
- AI agents
- Model evaluation
AI Evaluation
- Golden datasets
- Automated evaluation
- LLM-as-a-judge approaches
- Regression testing
- Evaluation pipelines
- Error classification
- Quality benchmarking
LLMOps / MLOps
- AI observability
- Model/version management
- Production monitoring
- CI/CD
- Release governance
- Experiment tracking
RAG / Context
- RAG
- Embeddings
- Vector databases
- Retrieval optimization
- Context engineering
- Memory architectures
Engineering
- Python
- REST/API integrations
- Cloud platforms
- Logging/monitoring
- Data pipelines
- Automated testing
Optimization
- Token optimization
- Prompt optimization
- Model routing
- Caching
- Latency optimization
- AI cost optimioptimisationzation
Production LLM/GenAI experience, clinical/healthcare AI experience, AI evaluation and regression, LLMOps/observability, and AI quality and optimisation.
📌 Senior AI/ML Engineer (Hyderabad)
🏢 Ekshvaku Tech Innovations
📍 Hyderabad