12 Sep
|
CriticalRiver
|
Hyderabad
12 Sep
CriticalRiver
Hyderabad
Job Summary
We are seeking a Senior AI / ML Architect with 10–15 years of experience and strong hands-on expertise in machine learning, deep learning, LLM model training, domain-specific fine-tuning, cloud deployment, and model hosting .
The primary focus of this role will be to train and fine-tune open-weight AI/ML models using domain-specific enterprise data and build scalable infrastructure to host, serve, manage, and monitor these models in production .
This is a highly hands-on role requiring strong practical experience with PyTorch, Hugging Face, LoRA, QLoRA, GPU infrastructure, cloud platforms, model serving, and MLOps tools . The candidate should be an experienced technical architect who can make architecture decisions while remaining deeply involved in implementation and troubleshooting.
Key Responsibilities
- Architect and implement solutions for training and fine-tuning ML and LLM models using domain-specific enterprise data .
- Evaluate and select appropriate open-weight models based on domain requirements, model capability, size, performance, licensing, infrastructure requirements, and cost.
- Build end-to-end pipelines for training data preparation, model training, fine-tuning, evaluation, model packaging, deployment, and monitoring .
- Perform hands-on LoRA and QLoRA fine-tuning of open-weight models.
- Configure and optimize fine-tuning parameters including rank, alpha, target modules, learning rate, batch size, sequence length, quantization, and training strategy .
- Work with domain-specific datasets to perform data cleaning, transformation, instruction formatting, dataset curation, synthetic data generation, deduplication, and quality validation .
- Train and fine-tune models using PyTorch, Hugging Face Transformers, PEFT, TRL, Accelerate , and related frameworks.
- Build scalable GPU-based training environments on AWS, Azure, or GCP.
- Plan and optimize GPU infrastructure for model training and inference, including A100, H100, L40S , or equivalent GPU environments.
- Deploy and host trained models using production-grade inference frameworks such as vLLM, SGLang, TensorRT-LLM, or TGI .
- Design scalable model-serving architectures supporting real-time inference, high availability, throughput, latency, and cost optimization .
- Manage the complete model lifecycle, including model versioning, registry, deployment, rollback, monitoring, upgrades, and retirement .
- Implement MLOps pipelines using tools such as MLflow, W&B;, Kubernetes, Docker, Ray, and CI/CD platforms .
- Monitor model and infrastructure performance, including model quality, latency, throughput, GPU utilization, memory utilization, failures, and cost .
- Implement model evaluation and regression frameworks to validate that fine-tuned models deliver measurable improvement over baseline models.
- Troubleshoot model training, GPU, inference, deployment, and production issues.
- Work closely with Data Scientists, ML Engineers, Data Engineers, and Cloud Engineers to operationalize models.
- Establish reusable architecture patterns and engineering standards for domain-specific model training and hosting .
- Provide technical leadership and mentor engineers working on ML/LLM training and deployment.
Required Skills 1. Model Training & Fine-Tuning — Non-Negotiable
- Strong hands-on experience with PyTorch and Hugging Face Transformers .
- Proven experience training and fine-tuning LLM/Deep Learning models using domain-specific data .
- Strong hands-on expertise in LoRA and QLoRA .
- Experience with PEFT, SFT, TRL, Accelerate , or equivalent frameworks.
- Experience with open-weight models such as Llama, Qwen, Mistral/Mixtral, Gemma, DeepSeek, or Phi .
- Strong understanding of training datasets, hyperparameters, GPU utilization, model evaluation, and training optimization.
2. Model Hosting & Cloud Deployment — Non-Negotiable
- Robust hands-on experience deploying and hosting ML/LLM models on AWS, Azure, or GCP .
- Experience with GPU-based cloud infrastructure and capacity planning.
- Hands-on experience with Docker and Kubernetes .
- Experience with model serving frameworks such as vLLM, SGLang, TensorRT-LLM, or TGI .
- Ability to design scalable and reliable model-serving infrastructure for production workloads.
3. MLOps & Model Lifecycle Management — Non-Negotiable
- Hands-on experience implementing ML/LLM lifecycle management from experimentation through production.
- Experience with MLflow, W&B;, model registries, CI/CD, model versioning, deployment, monitoring, and rollback .
- Ability to establish automated pipelines for training, evaluation, deployment, and retraining .
- Strong understanding of production model monitoring and operational management.
Experience & Qualifications
- 10–15 years of overall technology experience .
- Strong experience in ML/AI engineering, Deep Learning, and production model development.
📌 Senior AI / ML Architect – Model Training, Fine-Tuning & Hosting (Hyderabad)
🏢 CriticalRiver
📍 Hyderabad