BCI has an open position on our offshore GenAI team working with our USA based client. The GenAI / Python / LLM Fine Tuning Engineer will join our offshore development team that is growing and there is a lot of new and exciting GenAI work to be completed. This is a full-time remote position and must be able to work blended hours of EST / IST timings.
Please note: In addition to doing GenAI / Python Development, this role will also focus on LLM Fine tuning - LoRA / QLoRA, PEFT, full fine-tuning, instruction tuning, RLHF/DPO. Fine-tune and adapt large language models (Llama, Gemma, and other open-weight models.
About the Role
We're looking for an GenAI / Python / LLM Fine Tuning Engineer for the customization and optimization of large language models for production use cases. This role involves full fine-tuning lifecycle — from data preparation through training, evaluation, and deployment — working with open-weight models (e.g., Llama, Gemma) as well as proprietary/managed models (e.g., Google Gemini) where fine-tuning access is available. You'll be hands-on with real training runs at scale, not just prompt engineering, or API integration.
What You'll Do
- Fine-tune and adapt large language models (Llama, Gemma, and other open-weight models, plus managed options like Gemini where applicable) for specific business use cases
- Design and execute full fine-tuning pipelines: dataset curation and cleaning, tokenization, training/eval splits, hyperparameter selection, and training runs (full fine-tune, LoRA/QLoRA, PEFT, RLHF/DPO as appropriate)
- Run and manage large-scale training jobs across multi-GPU / distributed environments
- Evaluate model performance using both automated benchmarks and human-in-the-loop review; iterate to close quality gaps
- Optimize models for production inference (quantization, distillation, latency/cost tradeoffs)
- Deploy fine-tuned models into production systems and monitor performance, drift, and degradation over time
- Collaborate with data, ML infrastructure, and product teams to define fine-tuning objectives and success metrics
- Stay current on the open-model landscape and evaluate new base models as candidates for fine-tuning
- Document methodology, training runs, and results for reproducibility and knowledge sharing
Required Qualifications
- Strong proficiency in Python, with solid software engineering fundamentals (not just notebooks)
- Hands-on, production-level experience GenAI and fine-tuning LLMs — this is a must-have, not exploratory/academic experience only
- Demonstrated experience taking fine-tuned models into live, large-scale production systems (not just POCs)
- AWS and /or Google cloud production experience will be considered
- Experience with open-weight model families (e.g., Llama, Gemma, Mistral, or similar)
- Practical knowledge of fine-tuning techniques: LoRA/QLoRA, PEFT, full fine-tuning, instruction tuning, RLHF/DPO
- Experience with ML/training frameworks such as PyTorch,
Hugging Face Transformers/TRL/PEFT, DeepSpeed, or similar
- Familiarity with distributed/multi-GPU training and the associated infrastructure challenges
- Solid understanding of model evaluation methodology for generative models
Nice to Have
- Experience fine-tuning or customizing Google Gemini or other managed/API-based models
- Experience with vector databases, RAG architectures, or hybrid RAG + fine-tuning approaches
- Experience with MLOps tooling for training pipelines (e.g., MLflow, Weights & Biases, Kubeflow, SageMaker, Vertex AI)
- Experience with model quantization and inference optimization (vLLM, TensorRT-LLM, GGUF, etc.)
- Background in NLP research or publications related to LLM training/fine-tuning
What Success Looks Like Within your first few months, you're independently running fine-tuning jobs on open models, have a transparent point of view on which base models and techniques fit which use cases, and have shipped at least one fine-tuned model into a production system with measurable quality improvement over baseline.
Interview Process:
1. If profile appears to fit role, we will send you a request for more information and details on your background. 2. Initial 30 min MS Teams conversation with BCI-IT team to go over your hands on experience and determine fit. 3. If potential fit, you will be sent a video technical screen with several questions on LLM fine tuning. 4. 45 min to 1 hour client technical interview with code share activity and technical discussion. You will be speaking with 2-3 Sr. team members. Hiring decision can be made after call.
📌 Sr. GenAI / Python / LLM Fine Tuning Engineer (Delhi)
🏢 BCI~IT
📍 Delhi