11 Sep
|
CriticalRiver
|
Hyderabad
11 Sep
CriticalRiver
Hyderabad
Job Summary
We are seeking a Senior AI / ML Architect with 10–15 years of experience and strong hands-on expertise in machine learning, deep learning, LLM model training, domain-specific fine-tuning, cloud deployment, and model hosting.
The primary focus of this role will be to train and fine-tune open-weight AI/ML models using domain-specific enterprise data and build scalable infrastructure to host, serve, manage, and monitor these models in production.
This is a highly hands-on role requiring strong practical experience with PyTorch, Hugging Face, LoRA, QLoRA, GPU infrastructure, cloud platforms, model serving, and MLOps tools. The candidate should be an experienced technical architect who can make architecture decisions while remaining deeply involved in implementation and troubleshooting.
Key Responsibilities
- Architect and implement solutions for training and fine-tuning ML and LLM models using domain-specific enterprise data.
- Evaluate and select appropriate open-weight models based on domain requirements, model capability, size, performance, licensing, infrastructure requirements, and cost.
- Build end-to-end pipelines for training data preparation, model training, fine-tuning, evaluation, model packaging, deployment, and monitoring.
- Perform hands-on LoRA and QLoRA fine-tuning of open-weight models.
- Configure and optimize fine-tuning parameters including rank, alpha, target modules, learning rate, batch size, sequence length, quantization, and training strategy.
- Work with domain-specific datasets to perform data cleaning, transformation, instruction formatting, dataset curation, synthetic data generation, deduplication, and quality validation.
- Train and fine-tune models using PyTorch, Hugging Face Transformers, PEFT, TRL, Accelerate, and related frameworks.
- Build scalable GPU-based training environments on AWS, Azure, or GCP.
- Plan and optimize GPU infrastructure for model training and inference, including A100, H100, L40S, or equivalent GPU environments.
- Deploy and host trained models using production-grade inference frameworks such as vLLM, SGLang, TensorRT-LLM, or TGI.
- Design scalable model-serving architectures supporting real-time inference, high availability, throughput, latency, and cost optimization.
- Manage the complete model lifecycle, including model versioning, registry, deployment, rollback, monitoring, upgrades, and retirement.
- Implement MLOps pipelines using tools such as MLflow, W&B;, Kubernetes, Docker, Ray, and CI/CD platforms.
- Monitor model and infrastructure performance, including model quality, latency, throughput, GPU utilization, memory utilization, failures, and cost.
- Implement model evaluation and regression frameworks to validate that fine-tuned models deliver measurable improvement over baseline models.
- Troubleshoot model training, GPU, inference, deployment, and production issues.
- Work closely with Data Scientists, ML Engineers, Data Engineers, and Cloud Engineers to operationalize models.
- Establish reusable architecture patterns and engineering standards for domain-specific model training and hosting.
- Provide technical leadership and mentor engineers working on ML/LLM training and deployment.
Required Skills
1. Model Training & Fine-Tuning — Non-Negotiable
- Strong hands-on experience with PyTorch and Hugging Face Transformers.
- Proven experience training and fine-tuning LLM/Deep Learning models using domain-specific data.
- Robust hands-on expertise in LoRA and QLoRA.
- Experience with PEFT, SFT, TRL, Accelerate, or equivalent frameworks.
- Experience with open-weight models such as Llama, Qwen, Mistral/Mixtral, Gemma, DeepSeek, or Phi.
- Strong understanding of training datasets, hyperparameters, GPU utilization, model evaluation, and training optimization.
2. Model Hosting & Cloud Deployment — Non-Negotiable
- Strong hands-on experience deploying and hosting ML/LLM models on AWS, Azure, or GCP.
- Experience with GPU-based cloud infrastructure and capacity planning.
- Hands-on experience with Docker and Kubernetes.
- Experience with model serving frameworks such as vLLM, SGLang, TensorRT-LLM, or TGI.
- Ability to design scalable and reliable model-serving infrastructure for production workloads.
3. MLOps & Model Lifecycle Management — Non-Negotiable
- Hands-on experience implementing ML/LLM lifecycle management from experimentation through production.
- Experience with MLflow, W&B;, model registries, CI/CD, model versioning, deployment, monitoring, and rollback.
- Ability to establish automated pipelines for training, evaluation, deployment, and retraining.
- Strong understanding of production model monitoring and operational management.
Experience & Qualifications
- 10–15 years of overall technology experience.
- Strong experience in ML/AI engineering, Deep Learning, and production model development.
📌 Senior AI / ML Architect – Model Training, Fine-Tuning & Hosting (Hyderabad)
🏢 CriticalRiver
📍 Hyderabad