Generative AI Engineer (India)

Generative AI Engineer (India)

02 Aug
|
TOPS Infosolutions
|
India

02 Aug

TOPS Infosolutions

India

Gen AI Engineer LLM &

- Agentic Systems

Experience : 36 years total engineering, with 1.5 years shipping LLM-powered systems to production.

Location : Ahmedabad (Work from Office).

About the Role :

We're hiring an engineer who has actually put LLM systems in front of real users not demo notebooks, not a RAG prototype that worked once on a clean PDF. You'll own agentic features end-to-end : designing the agent loop, building the tools it calls, wiring up MCP servers, managing context, and then doing the unglamorous work of making it fast, cheap, and reliable enough to keep in production.

This is an engineering role first. The models are a component, not the job. You should be comfortable writing production TypeScript and/or Python, reasoning about latency and cost budgets, and defending your design choices to clients and to the engineers around you.

You'll work alongside our full-stack team, so the systems you build ship inside real applications with real users, deadlines, and support burden.

What You'll Do :

- Design and ship LLM-powered features into production : agentic workflows, chat systems, retrieval, extraction, and automation.
- Own the agent architecture planning and execution loops, tool design, state and memory, failure recovery, human-in-the-loop checkpoints.
- Build and integrate MCP servers and clients : tool and resource schema design, transports, auth, and protected exposure of internal systems to models.
- Build agentic chat experiences : streaming, interruption and resumption, multi-turn state, context window management, citations, and tool-call rendering.
- Make it cheap and fast prompt and context engineering, caching, model routing, batching, smaller models where they hold up, measured against real cost and latency budgets.
- Build the evaluation layer : test sets, regression suites, LLM-as-judge where appropriate, and honest measurement of whether a change actually helped.
- Instrument everything tracing, token and cost accounting, failure taxonomies, and production monitoring for non-deterministic systems.
- Design guardrails and failure handling : validation, retries, fallbacks, prompt injection defense, and sane behavior when the model is wrong.
- Participate in client conversations scoping what's actually feasible, setting expectations on accuracy, and saying no to the demo that won't survive contact with production.
- Review code, mentor engineers on LLM engineering practice,



and raise the team's judgment about where this technology helps versus where it's the wrong tool.

Must-Have :

- Real production LLM experience systems with live users, where you owned accuracy, cost, latency, and the on-call consequences.
- Strong software engineering fundamentals production TypeScript/Node and/or Python
- API design, async patterns, error handling, testing; you'd be a solid engineer even without the AI layer.
- Deep LLM integration provider APIs (OpenAI, Anthropic, Google, open-weight models), streaming, structured outputs, function/tool calling, multimodal inputs, rate limits, retries, and graceful degradation.
- Agentic systems, hands-on you've built and debugged agent loops, and can explain concretely where they break : runaway tool calls, context rot, silent failures, compounding errors across steps.
- Agentic chat implementation streaming UX, conversation state and persistence, memory and summarization strategy, context assembly, and tool-call/artifact rendering in a real interface.
- MCP practical experience building or integrating MCP servers and clients, with a view on tool granularity, schema design, and what belongs in a tool versus a prompt.
- RAG done properly chunking and indexing strategy, embeddings, hybrid and semantic search, reranking, and evaluation of retrieval quality separately from generation quality.
- Vector and traditional data stores pgvector, Pinecone, Qdrant, Weaviate or similar, plus SQL/NoSQL; able to pick the right store and design schemas that survive change.
- Prompt and context engineering as an engineering discipline versioned, tested, and measured, not tweaked until the demo passes.
- Evals and observability you measure quality before and after changes, and can show what you measured with tools like Langfuse, LangSmith, Braintrust, or your own harness.
- Cost and latency optimization token accounting, caching strategies, model selection and routing, and knowing when a smaller model or plain code beats a bigger model.
- Cloud deployment AWS (or GCP/Azure)



for running these workloads : queues, async workers, background jobs, secrets, and debugging production issues.
- Git workflows in a team branching, reviews, conflict resolution.
- Strong communication you can explain non-deterministic system behavior to a client without overpromising or hiding behind jargon.
- Independence you take an ambiguous problem, scope what's realistic, and ship it.

Strong Plus (one or more) :

1. Agent &

- Orchestration Depth :
- Multi-agent systems, delegation and handoff patterns, orchestration frameworks (LangGraph, Mastra, OpenAI Agents SDK, or your own).
- Sandboxed code execution, computer/browser use, long-running and resumable agents.
- Prompt injection and tool-abuse threat modeling.

2.

Model Work :

- Fine-tuning, LoRA/PEFT, distillation into smaller task-specific models.
- Self-hosting and serving open-weight models (vLLM, Ollama, Bedrock, SageMaker).
- Embedding model selection and evaluation, custom rerankers.

3. Product &

- Platform :
- Full-stack ability to ship the interface as well as the backend (React/Next.js).
- Streaming protocols and real-time systems (SSE, WebSockets).
- CI/CD, Docker, infrastructure-as-code, and eval pipelines that run in CI.

4.

Data :

- Document ingestion at scale : OCR, layout-aware parsing, messy real-world PDFs.
- Knowledge graphs, structured extraction, and schema-constrained generation.

Nice to Have :

- Speech/voice pipelines (STT, TTS, realtime APIs).
- Security fundamentals : auth flows, secrets management, data governance for model inputs.
- Automated testing depth beyond the AI layer (unit, integration, e2e).

What We're Not Looking For :

- Prompt engineers who don't ship code this is a software engineering role with an AI specialization, not the other way round.
- Framework tourists we care whether you understand the agent loop, not which orchestration library you're loyal to this quarter.
- Demo builders a Streamlit app that impressed a stakeholder isn't production experience.
- Resume keyword bingo listing every vector database and framework tells us less than one system you took to production and can talk about honestly, including what went wrong.
- Hype merchants we want someone who will tell a client when an LLM is the wrong solution.
- Engineers who can't or won't talk to clients communication is part of the job.

📌 Generative AI Engineer (India)
🏢 TOPS Infosolutions
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: generative ai engineer (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: generative ai engineer (india) / india