16 Aug
|
HuntingCube Recruitment Solution
|
India
16 Aug
HuntingCube Recruitment Solution
India
About the Role :
We are looking for a Machine Learning Engineer to build and deploy production-grade AI systems powered by Large Language Models (LLMs). In this role, you will work on retrieval-augmented generation (RAG), search, multilingual NLP, document understanding, evaluation frameworks, and AI-powered workflows. You will collaborate closely with research, engineering, and product teams to deliver reliable, scalable AI solutions used in real-world environments.
Key Responsibilities :
- Design, build, and deploy production-ready LLM and NLP applications.
- Develop retrieval and search pipelines using RAG, embeddings, vector databases, and reranking techniques.
- Build AI agents and workflow automation for complex business use cases.
- Fine-tune and adapt open-source LLMs using techniques such as LoRA, QLoRA, or PEFT.
- Develop multilingual NLP solutions including summarization, translation, information extraction, and question answering.
- Design evaluation frameworks to measure model quality, hallucinations, factuality, latency, and overall user experience.
- Build feedback pipelines that capture production failures and improve model performance.
- Work closely with data, engineering, and product teams to translate business requirements into scalable ML solutions.
- Deploy and maintain ML models in production while ensuring scalability, monitoring, and reliability.
- Continuously improve system performance, inference efficiency, and operational costs.
Required Skills :
- 3-6 years of experience building production Machine Learning or NLP systems.
- Strong programming skills in Python.
- Hands-on experience with PyTorch and Hugging Face Transformers.
- Experience working with Large Language Models (LLMs).
- Solid understanding of Retrieval-Augmented Generation (RAG)
architectures.
- Experience with vector search technologies such as FAISS, Pinecone, Milvus, Chroma, or similar.
- Experience with prompt engineering, LLM evaluation, and model fine-tuning.
- Familiarity with LangChain, LangGraph, LlamaIndex, or similar orchestration frameworks.
- Experience building REST APIs using FastAPI or similar frameworks.
- Experience deploying ML solutions using Docker and cloud platforms (AWS, Azure, or GCP).
- Good understanding of software engineering best practices, version control, testing, and CI/CD.
Preferred Skills :
- Experience with AI agents or multi-agent systems.
- Experience with multilingual NLP applications.
- Experience with recommendation systems or semantic search.
- Experience with vLLM, Text Generation Inference (TGI), or LLM serving frameworks.
- Familiarity with ML evaluation tools such as Arize Phoenix, LangSmith, or similar.
- Experience working with enterprise AI applications in domains such as legal, healthcare, finance, or public sector.
- Publications or open-source contributions in AI/ML are a plus.
Nice to Have :
- Experience with distributed training, DeepSpeed, or FSDP.
- Experience optimizing inference latency and deployment costs.
- Knowledge of model monitoring, observability, and experimentation frameworks.
- Familiarity with cloud-native architectures and microservices.
What You'll Work On :
- Production LLM applications
- Retrieval-Augmented Generation (RAG)
- AI Agents &
- Workflow Automation
- Search &
- Recommendation Systems
- Multilingual NLP
- Evaluation &
- Observability
- Document Intelligence
- Prompt Engineering &
- Fine-tuning
- Cloud-native ML Deployment
If you're passionate about building reliable AI systems that solve real-world problems at scale, we'd love to hear from you.
📌 Machine Learning Engineer - Speech Recognition (India)
🏢 HuntingCube Recruitment Solution
📍 India