17 Sep
|
TVS Supply Chain Solutions
|
Madurai
17 Sep
TVS Supply Chain Solutions
Madurai
Senior GenAI & Agentic AI Engineer
LLM Engineer | Agentic AI Engineer | AI/ML Engineer | NLP Engineer | MLOps Engineer
Organisation: TVS SCS (TVSSCS) | Location:Chennai,Madurai,Coimbatore | Experience: 3-7 Years | Type: Full-Time
Skills & Technology Keywords
This section is structured to support automated candidate matching on job portals and AI-powered ATS platforms.
AI / LLM Platforms & APIs
OpenAI GPT-4 / GPT-4o | Anthropic Claude | Google Gemini | Meta LLaMA 2 / LLaMA 3 | Mistral | Falcon | Cohere | Azure OpenAI Service | Google Vertex AI | AWS Bedrock | Hugging Face Hub | Ollama
Agentic AI & Orchestration Frameworks
LangChain | LlamaIndex | AutoGen | CrewAI | Semantic Kernel | AgentExecutor | ReAct Agents | Chain-of-Thought Prompting | Tool Use / Function Calling | Multi-Agent Systems | Autonomous Agents | Workflow Orchestration | Human-in-the-Loop (HITL) | LangGraph
RAG & Knowledge Retrieval
Retrieval-Augmented Generation (RAG) | Vector Search | Semantic Search | Hybrid Search | Re-ranking | Contextual Chunking | Document Embeddings | sentence-transformers | OpenAI Embeddings | BGE Embeddings | E5 Embeddings | FAISS | Pinecone | Weaviate | Milvus | Chroma | Qdrant | pgvector | Elasticsearch
ML / Deep Learning & Fine-Tuning
PyTorch | TensorFlow | Hugging Face Transformers | PEFT | LoRA | QLoRA | Supervised Fine-Tuning (SFT) | RLHF | DPO | Instruction Tuning | Model Quantization | ONNX | TensorRT | vLLM | TensorRT-LLM | Triton Inference Server | Distillation
MLOps & Model Lifecycle
MLflow | Weights & Biases (W&B;) | DVC | Kubeflow | Airflow | Prefect | Model Registry | Model Versioning | CI/CD for ML | GitHub Actions | GitLab CI | Jenkins | ArgoCD | Experiment Tracking | Data Versioning
Cloud & Infrastructure
AWS (SageMaker, EC2, S3, Lambda, ECS) | Azure (OpenAI, ML Studio, AKS) | GCP (Vertex AI, Cloud Run, GKE) | Docker | Kubernetes | Helm | Terraform | Serverless | GPU Infrastructure | NVIDIA CUDA | A100 / H100 GPUs
Backend & API Engineering
Python | FastAPI | Flask | Django | Node.js | REST API | GraphQL | Async Programming | Microservices | gRPC | WebSockets | API Gateway | OAuth2 | JWT | Rate Limiting
Databases & Storage
PostgreSQL | MySQL | Redis | MongoDB | Elasticsearch | pgvector | SQL | NoSQL | Schema Design | Query Optimization | Object Storage (S3/Blob)
Observability & Monitoring
Prometheus | Grafana | LangSmith | LangFuse | Arize AI | Evidently AI | OpenTelemetry | Structured Logging | Distributed Tracing | Cost Monitoring | Hallucination Detection | Token Usage Tracking | Agent Trace Logging
Prompt Engineering & AI Safety
Prompt Engineering | System Prompts | Few-Shot Learning | Zero-Shot Prompting | Context Window Management | Guardrails | Responsible AI | AI Governance | Bias Detection | Output Validation | Grounding | Factuality
About the Role
TVSSCS is hiring a Senior GenAI & Agentic AI Engineer who has shipped production AI systems to real users - not just built prototypes. You will own the architecture, deployment, and reliability of LLM-powered and agentic AI systems that run in production at scale within the TVSSCS engineering organisation. This is an engineering-first role; the expectation is production-grade code, observable systems,
and measurable business impact.
Key Responsibilities
1. Agentic AI System Design & Deployment
· Architect and deploy multi-agent orchestration pipelines using LangChain, LlamaIndex, AutoGen, CrewAI, LangGraph, or Semantic Kernel
· Build autonomous AI agents with tool-use, self-correction, ReAct loops, and long-horizon task execution
· Implement short-term, long-term, and episodic memory layers for stateful, context-aware agent behavior
· Design human-in-the-loop (HITL) workflows for high-stakes decision points within agent pipelines
· Integrate function calling and structured output enforcement with GPT-4, Claude, Gemini, and open-source LLMs
2. LLM Integration, RAG & Knowledge Systems
· Build and optimize production RAG pipelines using Pinecone, Weaviate, Milvus, Chroma, Qdrant, or pgvector
· Apply advanced retrieval: hybrid search (dense + sparse), cross-encoder re-ranking, contextual chunking, HyDE
· Manage embedding models (OpenAI, BGE, E5, sentence-transformers) and vector index lifecycle in production
· Implement multi-LLM routing (OpenAI, Anthropic Claude, Azure OpenAI, AWS Bedrock, Google Vertex AI) with fallback and cost controls
· Apply prompt engineering best practices: chain-of-thought, few-shot, system prompt design, context window management
3. Production ML Engineering & MLOps
· Own model deployment end-to-end: Docker, Kubernetes, CI/CD (GitHub Actions / GitLab CI / ArgoCD), rollback strategies
· Optimize LLM inference for low-latency, high-throughput serving using vLLM, TensorRT-LLM, Triton Inference Server, ONNX
· Manage ML lifecycle with MLflow, Weights & Biases, DVC - experiment tracking, model registry, versioning
· Execute efficient fine-tuning with LoRA, QLoRA (PEFT), SFT, and RLHF / DPO on proprietary datasets
· Deploy on AWS (SageMaker, ECS, Lambda), Azure (ML Studio, AKS), or GCP (Vertex AI, Cloud Run)
4. Observability, Monitoring & Reliability
· Build real-time GenAI monitoring dashboards using LangSmith, LangFuse, Arize AI, Evidently AI, or Prometheus + Grafana
· Track production GenAI metrics: token costs, hallucination rates, latency p50/p95/p99, agent tool-call traces
· Implement OpenTelemetry-based distributed tracing and structured logging for multi-agent pipelines
· Design alerting, circuit breakers, and graceful degradation patterns for AI service failures
· Conduct root-cause analysis on production incidents and drive SLA/SLO improvements
5. Backend Engineering & API Development
· Build scalable AI microservices with FastAPI, Flask, or Django; design REST and WebSocket APIs for real-time AI features
· Implement authentication (OAuth2, JWT), rate limiting, input validation, and structured error handling
· Integrate with PostgreSQL, Redis, MongoDB, Elasticsearch for data persistence and caching layers
· Apply async programming patterns (asyncio, Celery) for high-concurrency AI workloads
Must-Have Requirements
· 3-7 years of hands-on experience in ML Engineering, AI Engineering, or LLM Engineering roles
· MANDATORY: At least 2 production-deployed AI/ML systems (describe in application: problem, scale, your ownership, tech stack used)
· Proven experience deploying LLM applications - RAG pipelines, AI chatbots, autonomous agents, or AI-powered workflows - to production cloud (AWS / Azure / GCP)
· Hands-on experience with at least one agentic AI framework: LangChain, LlamaIndex, AutoGen, CrewAI, or Semantic Kernel
· Strong Python proficiency; experience with FastAPI or Flask for building AI-serving microservices
· Experience integrating at least one LLM API in production: OpenAI, Anthropic, Azure OpenAI, Cohere, or Hugging Face
· Hands-on experience with vector databases (Pinecone / Weaviate / Milvus / Chroma / pgvector / FAISS) in production RAG systems
· Solid command of Docker and Kubernetes for containerized AI model deployment
· Understanding of prompt engineering, context window management, and LLM cost optimization
· Exposure to MLOps tooling: MLflow, Weights & Biases, or DVC for experiment tracking and model management
Good to Have
· Experience with multi-agent frameworks: AutoGen, CrewAI, LangGraph, or custom orchestrators
· LLM fine-tuning using LoRA, QLoRA, PEFT, SFT, RLHF, or DPO on domain-specific datasets
· Inference optimization experience: vLLM, TensorRT-LLM, ONNX, model quantization (INT4/INT8), or distillation
· GenAI observability experience with LangSmith, LangFuse, Arize AI, or Evidently AI
· Knowledge of responsible AI, AI governance, output guardrails, bias detection, and factuality validation
· Open-source contributions to GenAI, LLM, or agentic AI tooling
· Frontend exposure (React, TypeScript) for building internal AI dashboards or chat interfaces
· Familiarity with enterprise AI compliance, data privacy, and security best practices (SOC 2, GDPR)
What We Expect From Applicants
Candidates must demonstrate real production experience. In your application, include:
· Production system #1 - problem solved, users/requests at scale, your specific role, key technologies used
· Production system #2 - same format as above
· One production incident you debugged and resolved - what broke, how you diagnosed it, what you fixed
· Links to GitHub, portfolio, blog posts, or open-source contributions (preferred but not mandatory)
Note : Applications without documented production AI project experience will not be shortlisted.
Why Join TVSSCS
· Build enterprise-grade AI systems used in real business operations - not demos or internal tools
· Work with state-of-the-art LLM infrastructure: GPT-4, Claude, Gemini, LLaMA, custom fine-tuned models
· High ownership: you own architecture decisions, not just feature tickets
· Learning & certification budget for AI/ML conferences, cloud certifications, and research
· Collaborative engineering culture with solid peer review, design discussions, and knowledge sharing
- · Growth path into AI Architecture / Principal Engineer roles within TVSSCS latforms or
📌 GEN AI Engineers (Madurai)
🏢 TVS Supply Chain Solutions
📍 Madurai