08 Sep
|
AjnaLens
|
Thane
Namaskaram!
AjnaLens is looking for an AI/ML Engineer to join our Product Engineering team at Thane (Maharashtra – India). The ideal candidate should have 5+ years of experience building, integrating, optimizing, and deploying AI systems in real-world production environments. The role focuses on applying modern AI/ML techniques across Generative AI, LLMs, Computer Vision, multimodal AI, inference optimization, AI agents, and edge AI.
Candidates should be comfortable taking AI solutions from experimentation and prototyping through production deployment, monitoring, and optimization. This role demands a strong product-focused mindset, practical engineering skills, and the ability to translate cutting-edge AI capabilities into reliable, scalable, and efficient products.
We’re proud to share that Lenskart is now our strategic investor, a milestone that reflects the impact, potential, and purpose of the path we’re walking. Join us as we co-create the future of conscious, AI-powered technology.
? Read more here: The smartphone era is peaking. The next computing revolution is here.
Who are we looking for:
We are looking for a highly skilled AI/ML Engineer with strong hands-on experience in building, integrating, evaluating, and deploying AI solutions across Generative AI, LLMs, Computer Vision, multimodal AI, and AI agents. The candidate should have a solid foundation in machine learning and deep learning concepts, including supervised and unsupervised learning, model evaluation, neural network architectures, optimization, and transfer learning.
The ideal candidate should understand the complete AI lifecycle—from data preparation and experimentation to model training/fine-tuning, evaluation, model integration, inference optimization, API/model serving, deployment, monitoring, and production support.
Experience with LLM fine-tuning, RAG, prompt engineering, model quantization, inference optimization, and deploying AI workloads on GPUs/VMs is highly valuable.
Strong
Python skills and working knowledge of C++ for performance-critical applications are expected.
Top 3 Daily Tasks:
- Build, integrate, fine-tune, and optimize AI models and AI-powered features across Generative AI, LLM, Computer Vision, multimodal, and agentic AI use cases.
- Develop and deploy production-grade AI services, inference pipelines, and model-serving APIs using technologies such as FastAPI, Docker,
GPU runtimes, and cloud/on-premise VMs.
- Evaluate, monitor, troubleshoot, and optimize AI systems in production, focusing on latency, throughput, resource utilization, reliability, model quality, and cost.
Minimum work experience is required:
Minimum 5 years of hands-on experience in AI/ML engineering, production AI systems, model integration/deployment, and the end-to-end lifecycle of AI-powered applications.
Top 5 Skills you should possess:
- Strong proficiency in Python and solid working knowledge of C++ for AI integration, performance-critical systems, and inference optimization
- Robust understanding of Machine Learning and Deep Learning fundamentals, including model selection, feature engineering, supervised/unsupervised learning, neural network architectures, CNNs, transformers, transfer learning, training/validation, and model evaluation; hands-on experience with PyTorch or similar frameworks
- Hands-on experience with Generative AI, LLMs, LLM fine-tuning, prompt engineering, RAG, embeddings, vector databases, AI agents, and/or multimodal AI
- Practical experience with AI inference optimization, including quantization, model compression, batching, GPU utilization, latency/throughput optimization, and model serving
- Strong understanding of AI deployment, MLOps, model monitoring, API design, Docker, Linux, GPU-based infrastructure, and deploying AI workloads on VMs/cloud environments
Preferred / Good to Have:
- Experience with classical ML techniques such as regression, classification, clustering, feature engineering, ensemble methods, and model evaluation.
- Practical experience with deep learning architectures such as CNNs, RNNs/LSTMs, transformers, vision transformers, and transfer learning.
- Experience with training and fine-tuning models, hyperparameter optimization, experiment tracking, dataset versioning, and reproducible ML workflows.
What would you be expected to do:
- Experience deploying and managing AI inference workloads on Linux VMs, GPU VMs,
cloud infrastructure, or on-premise servers.
- Familiarity with inference runtimes and optimization stacks such as ONNX Runtime, TensorRT, vLLM, LiteRT/LiteRT-LM, or similar technologies.
- Experience with multimodal AI, vision-language models, speech/audio AI, or AI systems integrated with hardware products.
- Working knowledge of Kubernetes, CI/CD, Docker, observability, and production infrastructure for AI workloads.
- Experience with C++, CUDA, GPU profiling, or hardware-aware optimization is an added advantage.
What would you be expected to do:
- Design and implement end-to-end AI solutions that balance model quality, latency, scalability, reliability, infrastructure cost, and user experience.
- Integrate, fine-tune, evaluate, and deploy LLMs and Generative AI models for use cases such as conversational AI, summarization, entity extraction, classification, reasoning, text generation, and AI agents.
- Work with Computer Vision and multimodal AI systems, including image understanding, visual-language models, embeddings, classification, detection, and other perception workloads where applicable.
- Build and maintain REST APIs and model-serving services using FastAPI and related serving frameworks, with proper validation, documentation, logging, and observability.
- Create reproducible AI pipelines covering data preparation, experimentation, evaluation, model integration, deployment, and production validation.
- Evaluate AI systems using relevant quality and performance metrics, test edge cases, analyze failure modes, and establish reliable validation and benchmarking processes.
- Design prompt engineering strategies, RAG pipelines, guardrails, tool calling, and safety mechanisms for production Generative AI and agentic AI systems.
- Build and maintain AI observability and monitoring to track model quality, latency, resource utilization, failures, drift, anomalies, and out-of-distribution inputs.
- Optimize AI inference for low latency and high throughput using techniques such as quantization, batching, caching, efficient model serving, GPU optimization, and hardware-aware deployment.
- Deploy and operate AI workloads on Linux-based VMs and cloud/on-premise GPU infrastructure; manage runtime dependencies, containers, model artifacts, services, logs, and production troubleshooting.
📌 AI Engineer (Thane)
🏢 AjnaLens
📍 Thane