Location: Chennai — On-site
Employment Type: Full-time
Working Days: 6 Days a Week
Industry: SaaS / Product / AI Technology
Hiring Partner: AP Digital Innovations
About the Role
AP Digital Innovations is hiring on behalf of one of our clients in the SaaS/Product technology space for an AI / Generative AI Engineer.
This is a highly hands-on engineering role focused on building production-grade AI-powered applications, intelligent document processing systems, RAG pipelines, LLM-based workflows, and AI-assisted product experiences.
The ideal candidate should have strong experience with Python, PostgreSQL, LLMs, RAG, embeddings, vector databases, token management, prompt engineering, and AI application architecture.
You will work closely with product and engineering teams to design, build, integrate, and optimize AI capabilities for real-world SaaS applications.
Important: AP Digital Innovations is acting as a hiring partner for this opportunity. The client company name will be shared only with shortlisted candidates during the interview process.Key ResponsibilitiesAI / LLM Application Development
- Design and develop production-ready AI-powered applications using Python.
- Integrate commercial and open-source LLMs into SaaS products.
- Build intelligent AI workflows for document processing, classification, extraction, summarization, question answering, and decision support.
- Develop AI services and APIs that can be integrated with existing web applications.
- Evaluate and select appropriate LLMs based on accuracy, latency, cost, context size, and use case requirements.
- Work with both cloud-hosted and locally hosted/open-source AI models.
RAG & Knowledge Systems
- Design and implement Retrieval-Augmented Generation (RAG) pipelines.
- Build document ingestion, parsing, chunking, embedding, indexing, retrieval, and generation workflows.
- Develop semantic and hybrid search capabilities.
- Implement metadata filtering and context-aware retrieval.
- Optimize chunking strategies, retrieval quality, reranking, and context construction.
- Build systems that provide reliable answers grounded in enterprise documents and application data.
- Implement mechanisms to reduce hallucinations and improve response accuracy.
Embeddings & Vector Search
- Work with embedding models for semantic search and document similarity.
- Design and maintain vector search infrastructure.
- Experience with technologies such as pgvector, PostgreSQL, Qdrant, Pinecone, Weaviate, Milvus, or similar.
- Implement document-to-vector and query-to-vector pipelines.
- Optimize vector indexing, similarity search, filtering, and retrieval performance.
- Understand the trade-offs between dense retrieval, sparse retrieval, and hybrid search.
Token Management & LLM Optimization
- Implement effective token management strategies for LLM applications.
- Monitor input/output token usage and context-window limitations.
- Optimize prompts and retrieved context to reduce unnecessary token consumption.
- Implement conversation history management and context compression.
- Design strategies for handling long documents and large conversations.
- Optimize AI applications for latency, throughput, reliability, and API cost.
- Implement model fallback and routing strategies where required.
Prompt Engineering & AI Workflows
- Develop and maintain structured prompts for production AI workflows.
- Work with system prompts, few-shot prompting, structured outputs, and function/tool calling.
- Build multi-step AI workflows and agent-like systems where appropriate.
- Implement deterministic validation around LLM-generated outputs.
- Design JSON/schema-based AI responses for reliable application integration.
- Continuously evaluate and improve prompts based on real-world results.
Intelligent Document Processing
- Build AI pipelines for processing resumes, PDFs, documents, forms,
and other unstructured data.
- Implement document extraction and normalization workflows.
- Extract structured information from unstructured documents using LLMs and traditional processing techniques.
- Handle OCR-generated and imperfect document data.
- Build document classification, entity extraction, summarization, and enrichment pipelines.
- Design scalable asynchronous processing workflows for large document volumes.
Backend & API Engineering
- Develop scalable backend services using Python.
- Build REST APIs and AI service endpoints.
- Work with frameworks such as FastAPI, Flask, or Django.
- Integrate AI services with existing frontend and backend systems.
- Implement authentication, authorization, validation, logging, error handling, and monitoring.
- Design asynchronous and background processing systems for long-running AI jobs.
Database & Data Engineering
- Work extensively with PostgreSQL and relational data models.
- Design schemas for AI-related data and application workflows.
- Work with JSON/JSONB, indexes, queries, transactions, and performance optimization.
- Integrate PostgreSQL with vector search using pgvector or equivalent technologies.
- Build pipelines connecting structured application data with unstructured document knowledge.
AI Performance & Production Optimization
- Profile and optimize AI pipelines for production workloads.
- Improve inference latency and throughput.
- Evaluate model quality, retrieval quality, and end-to-end response quality.
- Implement caching where appropriate.
- Design retry, timeout, rate-limit, and fallback mechanisms.
- Monitor LLM usage, token consumption, latency, errors, and cost.
- Build reliable AI systems that can operate at scale.
AI Evaluation & Quality
- Create evaluation datasets and test cases for AI features.
- Measure retrieval accuracy, answer relevance, hallucination rate, and response quality.
- Compare different models, prompts, embeddings, and retrieval strategies.
- Establish regression testing for AI workflows.
- Implement guardrails and validation mechanisms for production AI systems.
Mandatory Requirements
- 4+ years of professional software engineering experience, with significant hands-on experience in AI/ML/GenAI application development.
- Strong proficiency in Python.
- Strong experience building production applications using LLMs.
- Strong understanding of RAG architecture.
- Hands-on experience with embeddings and vector databases/vector search.
- Strong experience with PostgreSQL.
- Understanding of tokenization, context windows, token usage, and LLM cost optimization.
- Strong knowledge of prompt engineering and structured LLM outputs.
- Experience building REST APIs using Python frameworks such as FastAPI or Flask.
- Understanding of document ingestion and processing pipelines.
- Experience integrating AI models into real-world SaaS/product applications.
- Strong understanding of software engineering principles, debugging, testing, and production deployment.
- Ability to design scalable and maintainable AI systems.
Preferred Skills
- Python
- FastAPI
- PostgreSQL
- pgvector
- RAG
- LLM APIs
- OpenAI-compatible APIs
- Anthropic / Gemini / other commercial LLM APIs
- Ollama
- Open-source LLMs
- Hugging Face
- Embedding models
- Vector databases
- Prompt engineering
- Function calling / tool calling
- Structured outputs
- AI agents and workflow orchestration
- Document AI
- Semantic search
- Hybrid search
- Reranking
- Redis
- Celery / background workers
- REST APIs
- Docker
- Git
- AWS
Good to Have
- Experience with LangGraph, LangChain, LlamaIndex, or similar AI orchestration frameworks.
- Experience deploying open-source LLMs.
- Experience with Ollama, vLLM, or similar inference platforms.
- Knowledge of transformer architectures and attention mechanisms.
- Understanding of model quantization and inference optimization.
- Experience with GPU-based inference.
- Experience building AI agents and tool-using systems.
- Experience with OCR and multimodal AI.
- Experience with speech, vision, or multimodal models.
- Knowledge of ML fundamentals and Python ML libraries.
- Experience with AWS services such as EC2, S3, Lambda, ECS, or related infrastructure.
- Experience building AI observability and evaluation systems.
- Experience working with high-volume enterprise documents.
Indicative Tech Stack
Programming:
Python
Backend:
FastAPI, Flask, REST APIs
Database:
PostgreSQL, pgvector, Redis
AI / LLM:
OpenAI-compatible APIs, Gemini, Anthropic, Ollama, Open-source LLMs
RAG:
Document ingestion, chunking, embeddings, vector search, hybrid retrieval, reranking, contextual generation
AI Frameworks:
LangGraph, LangChain, LlamaIndex or equivalent
Embeddings / Vector Search:
pgvector, Qdrant, Pinecone, Weaviate, Milvus or equivalent
Document Processing:
PDF processing, OCR, document parsing, structured extraction
Infrastructure:
AWS, Docker, CDN / Cloud services
Development:
Git, GitHub, CI/CD, automated testing, API testing
Experience Range
4–9 years
Salary Range
₹20–30 LPA
Work Mode
Work from Office — Bengaluru
Ideal Experience Level
Mid-level to Senior AI / GenAI Engineers
Ideal Candidate Profile
We are looking for an engineer who:
- Has strong hands-on Python development experience.
- Understands how up-to-date LLM applications actually work beyond simple API integration.
- Can independently design and implement a complete RAG pipeline.
- Understands embeddings, vector search, retrieval, context construction, and generation.
- Understands token limits and token economics and can optimize AI applications accordingly.
- Can work with both structured PostgreSQL data and unstructured documents.
- Has experience taking AI prototypes into reliable production systems.
- Understands that LLM output needs validation, evaluation, monitoring, and guardrails.
- Has strong backend engineering fundamentals.
- Can troubleshoot performance, reliability, model quality, and integration issues.
- Is comfortable working in a fast-paced SaaS/product environment.
- Has strong ownership and problem-solving skills.
- Is genuinely interested in building practical AI products rather than only experimenting with models.
Why Join
- Work on production-grade AI and GenAI applications.
- Build advanced RAG and intelligent document-processing systems.
- Work with modern LLMs, embeddings, vector search, and AI workflows.
- Solve challenging problems involving accuracy, context, latency, scalability, and AI cost.
- Work on SaaS products with real-world business use cases.
- Opportunity to work across the complete AI application lifecycle — from data ingestion and retrieval to LLM orchestration and production deployment.
- High-ownership role in a fast-growing technology environment.
How to Apply
Interested candidates can apply or share their updated resume at:
[email protected]
Key Note: AP Digital Innovations is acting as the hiring partner for this opportunity on behalf of its client. Client company details will be shared with shortlisted candidates during the interview process. Kindly avoid mentioning or requesting the client company name publicly on job portals or social media platforms.
Pay: ₹1,000,000.00 - ₹2,200,000.00 per year
Work Location: In person
📌 AI / Generative AI Engineer — Python, RAG & LLM (Chennai)
🏢 Ap Digital Innovations
📍 Chennai