About the Role
We re looking for an AI Engineer to design, build, and ship production-grade AI systems. You ll work across the full lifecycle - from experimenting with models and building retrieval pipelines to deploying and monitoring systems that serve real users at scale. This role sits at the intersection of applied machine learning and software engineering, so you ll need to be equally comfortable fine-tuning a model and writing the service that wraps it.
You ll partner closely with product, data, and platform teams to turn ambiguous problems into reliable, measurable AI features.
What You ll Do
- Build and productionize AI/ML systems, including large language model (LLM) applications, retrieval-augmented generation (RAG) pipelines, and model-serving infrastructure.
- Design and implement RAG architectures: chunking and embedding strategies, vector store selection and tuning, retrieval and re-ranking, and grounding responses to reduce hallucination.
- Fine-tune, adapt, and evaluate models using PyTorch, TensorFlow, and the Hugging Face ecosystem (Transformers, Datasets, Tokenizers, PEFT/LoRA).
- Develop robust evaluation frameworks - offline benchmarks, human-in-the-loop review, and online metrics - to measure quality, latency, and cost.
- Optimize inference for performance and cost (quantization, batching, caching, GPU utilization).
- Write clean, well-tested, maintainable code and contribute to shared libraries and services.
- Collaborate with product and design to scope AI features, and communicate trade-offs clearly to technical and non-technical stakeholders.
- Monitor deployed systems for drift, regressions, and reliability, and iterate based on real-world feedback.
Required Qualifications
- Bachelor s or Master s degree in Computer Science, Machine Learning, a related field, or equivalent practical experience.
- AI Frameworks:
Working knowledge of PyTorch, TensorFlow, and the Hugging Face ecosystem.
- RAG: Hands-on experience building retrieval-augmented generation systems - embeddings, vector databases (e.g., FAISS, Pinecone, Weaviate, pgvector), and retrieval pipeline design.
- LLMs: Practical experience working with large language models, including prompting, fine-tuning, and evaluation.
- Strong programming skills in Python, with solid software engineering fundamentals (version control, testing, code review).
- Experience deploying models or ML services to production (APIs, containers, cloud environments).
- Familiarity with the ML lifecycle: data preparation, experimentation, evaluation, deployment, and monitoring.
- Ability to read research papers and translate ideas into working implementations.
Preferred Qualifications
- Experience with orchestration frameworks such as LangChain, LlamaIndex, or equivalent.
- Familiarity with agentic systems, tool use, and function calling.
- Experience with inference optimization (vLLM, TensorRT, ONNX, quantization) and GPU workloads.
- Exposure to MLOps tooling (MLflow, Weights & Biases, Kubeflow) and CI/CD for ML.
- Cloud experience (AWS, GCP, or Azure) and infrastructure-as-code.
- Knowledge of responsible AI practices: evaluation for safety, bias, and reliability.
- Open-source contributions or a portfolio of applied AI projects.
What We Offer
- Competitive salary and equity.
- [Health, retirement, and other advantages].
- Budget for compute, learning, and conferences.
- A collaborative team building AI products that reach real users.
- [Flexible work arrangements / other perks].
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 AI Engineer (Gurugram)
🏢 Wiom
📍 Gurugram