27 Aug
|
Stefanini Group
|
Mumbai
27 Aug
Stefanini Group
Mumbai
Details:
We are seeking an experienced AI Engineer with 3-6 years of hands-on experience designing, developing, and deploying generative AI applications in production environments. The candidate will be responsible for building intelligent, AI-powered features - including text generation, summarization, conversational AI, and agentic workflows - and integrating them securely into scalable, cloud-based backend systems.
The role requires a solid foundation in large language model (LLM) systems, including prompt engineering, retrieval-augmented generation (RAG), agent orchestration, and output evaluation, combined with solid backend development expertise. Experience with the Google AI ecosystem (Gemini API, Vertex AI, Agent Development Kit) is an advantage; candidates with equivalent experience on other major LLM platforms are encouraged to apply.
- Design, develop, and deploy generative AI features such as text generation, summarization, conversational assistants, and multi-step agentic workflows.
- Architect and implement retrieval-augmented generation (RAG) pipelines, covering document ingestion, chunking, embeddings, vector store integration, retrieval and reranking, and grounding quality assessment.
- Develop agentic systems using tool/function calling, structured outputs, and orchestration patterns, incorporating appropriate guardrails, fallback mechanisms, and human-in-the-loop controls.
- Establish and maintain prompt engineering standards, including prompt versioning, structured output schemas, and data-driven optimization of response quality and accuracy.
- Build evaluation frameworks for LLM outputs, including curated test datasets, automated evaluations, regression testing, and monitoring for hallucination and grounding quality.
- Integrate AI services into backend applications through well-designed REST APIs and microservices, with robust handling of structured JSON responses, streaming, retries, and error states.
- Implement secure API authentication and access management for AI services, including API key management, OAuth 2.0, IAM, secrets handling, and safeguards against prompt injection and data leakage.
- Monitor and optimize production performance across response latency, token cost, throughput, and output quality, supported by appropriate observability and tracing.
- Collaborate with product managers, data engineers, and application developers to embed AI capabilities into business applications while ensuring security, reliability, and compliance.
Details:
- Bachelor's or Master's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 3-6 years of experience in AI/software engineering, including hands-on delivery of LLM-powered applications in production.
- Strong understanding of LLM fundamentals, including transformer architecture, tokenization, context windows, embeddings, sampling parameters, and the trade-offs between prompting, RAG, and fine-tuning.
- Working knowledge of common LLM failure modes (e.g., hallucination, prompt sensitivity, context degradation) and corresponding mitigation strategies.
- Hands-on experience with one or more major LLM platforms, such as Google Gemini, OpenAI, Anthropic Claude,
or open-source models (Hugging Face, vLLM).
- Practical experience building RAG systems, including chunking strategies, embedding models, vector databases (e.g., pgvector, Pinecone, Weaviate, Vertex AI Vector Search), and retrieval evaluation.
- Experience implementing tool/function calling and agentic workflows, using frameworks such as LangGraph, LangChain, Google ADK, or CrewAI, or through custom implementations.
- Proficiency in prompt engineering, supported by structured evaluation of output quality.
- Strong programming skills in Python; experience with Node.js or similar backend technologies is a plus.
- Solid backend engineering fundamentals, including REST API design, microservices architecture, and scalable, fault-tolerant system design.
- Experience integrating AI services within cloud architectures (GCP, AWS, or Azure), including secure API authentication and structured JSON response handling.
Preferred Qualifications
- Direct experience with the Google AI ecosystem, including Gemini API, Vertex AI, Google Agent Development Kit (ADK), or Google Antigravity.
- Experience with the Model Context Protocol (MCP) or building tool integrations for agentic systems.
- Exposure to model fine-tuning (e.g., LoRA/PEFT, instruction tuning) and model serving.
- Familiarity with LLM observability and evaluation tooling (e.g., LangSmith, Langfuse, Vertex AI Evaluation).
- Experience with modern front-end frameworks (React, Angular, or Vue.js) for building responsive, AI-driven user interfaces.
- Experience with containerization and orchestration (Docker, Kubernetes) and CI/CD practices.
- Familiarity with responsible AI practices, including content safety, PII handling, and compliance considerations.
📌 AI Engineer � Generative AI (Mumbai)
🏢 Stefanini Group
📍 Mumbai