27 Sep
|
Mondee
|
Hyderabad
Product Engineering Lead — AI & Agent Platform
Tabhi Inc. — Product Engineering —
Job Title: Product Engineering Lead — AI & Agent Platform
Posting Title: AI Architect — Agent Platform
Experience Required
8–9 Years
Location Hyderabad
Openings - 1–2
Reports To: VP, Product Engineering
Function: Product Engineering — works across multiple platforms
Work Location: Hyderabad
Work Mode: Work From Office
ABOUT THE ROLE
You will join the product engineering group and work across several AI-native platforms, not a single product . The common thread is the agent layer: assistants, deep agents and workforce agents, the tools and retrieval behind them, and the evaluation discipline that makes their behaviour something we can trust.
This is where correctness matters most. It is where a system can be confident, well-formatted and wrong — and in our domain, wrong has consequences for a real customer and a real transaction.
You report to the VP of Product Engineering, who owns platform-wide architecture and the roadmap. Inside the agent platform, the design and the delivery are yours: you turn direction into specifications, lead a small pod through them, write a good share of the code yourself, and own what ships.
Our stack: Next.js and React Native on the front end; Java and Spring Boot services; Python, FastAPI and LangGraph in the AI layer; MongoDB Atlas with vector search, Redis, Kafka and Elasticsearch; AWS and Oracle Cloud.
This is a hands-on lead role, not a management one. You will be in the codebase most days, working alongside code models the whole time.
HOW WE BUILD
We build with AI coding agents every day — Claude Code most of all. It is not autocomplete here; it is how the platform gets written, and it has changed what a senior engineer's day looks like. We are looking for someone already working this way.
- You have built several substantial applications end to end with Claude or comparable code models — not a few assisted commits.
- You structure repositories so an agent inherits the architecture rather than inventing one: context files, conventions, custom commands, sub-agents, hooks, MCP servers wired into your own loop.
- You specify before you generate, and you are clear about which decisions the agent does not get to make.
- You verify what comes back — tests before acceptance, real diff review, gates in CI. The failure mode of this way of working is merging something plausible and wrong.
- You keep exploring. New models and tools get a real evaluation when they land, and you share what you learn with the people around you.
We would genuinely like to see what you have built — repositories, context files, write-ups, or a walkthrough of a project you are proud of.
WHAT YOU WILL OWN
Agent architecture and orchestration
- Agent platforms across our products — assistants, deep agents and workforce agents, and the surfaces they appear on.
- Orchestration in LangGraph: graph design, state, retries, timeouts and degradation behaviour.
- Tool contracts — typed, versioned, and honest about failure. A tool that cannot answer must say so rather than fabricate.
- Human-in-the-loop and approval gates that genuinely halt execution, with a test proving it.
Architecture and technology choices
- Reference architectures and reusable patterns, so the tenth agent we build costs less than the third.
- Build-versus-buy across LLM providers, vector stores, orchestration frameworks and voice vendors, weighed on capability, cost, latency and lock-in.
- Choosing the right technique for each problem — retrieval, a tool call, a fine-tune, or a deterministic rule.
- Integration with the rest of the platform through versioned APIs and typed contracts.
Retrieval, memory and grounding
- Memory tiers: what is remembered, for how long, at what scope, and what must never be retained.
- Retrieval and vector strategy on MongoDB Atlas — chunking, hybrid search, reranking, freshness.
- Grounding and citation, so assertions the system makes can be traced to a source.
Evaluation and quality
- Evaluation suites and behavioural regression testing that run in CI, not in a notebook.
- Golden datasets drawn from real user interactions, grown from production failures.
- Quality gates for agent changes: what must pass before a graph change can merge.
- Cost per assisted session and p95 latency as owned, budgeted, alerted numbers.
Responsible AI and guardrails
- What the agent may assert, what it must cite, what it must refuse, and where it hands off to a human.
- Traceability and auditability of agent actions.
- Personal data in prompts, logs and traces — handling, retention and GDPR-aligned practice.
Delivery leadership
- Turn direction into decision-complete specifications your engineers and their agents can execute without guessing.
- Run the pod's cadence — scope, daily unblocking, dependencies with the other pods, and honest escalation early when a date is at risk.
- Own the review queue for your area. Reviewing agent-generated code carefully is the core of the job, not an interruption to it.
- Take proofs of concept through to production, including the hardening a validated prototype still needs.
- Own release readiness and report delivery health weekly.
Technical leadership
- Own the AI-native engineering practice for your pod — the context files, the templates, the conventions, the review habits — and help raise it across the group.
- Run architecture and code reviews, and mentor the engineers around you.
- Partner with product managers and domain experts across products to turn business problems into solution designs.
THE DOMAIN Our products sit in travel technology — booking, distribution and servicing across several channels. You do not need to arrive as a travel expert, but you will need to become one, because most serious agent failures in this space are domain failures wearing an engineering costume. What matters:
- The booking lifecycle — search, shop, book, ticket, service — and why a price quoted ninety seconds ago may no longer be bookable.
- Supplier reality. Content arrives from GDS, NDC and direct connections with different latency and reliability. Timeouts and partial results are normal conditions, not exceptions.
- Segment differences. The same question has a different correct answer for an agency, a consumer, a partner channel and a corporate traveller — and an agent that blurs them is a compliance problem, not a bug.
- Breadth across products. What you build in one platform should be reusable in the next. Designing for that is part of the job.
Travel domain depth is a real advantage and we weight it. If you do not have it, tell us; we care more about how rapid you pick up an unfamiliar, high-consequence domain — you will be doing that repeatedly as you move between products.
MUST HAVE
- 8–9 years of engineering experience, with a solid backend or distributed-systems foundation established before the LLM work.
- 2+ years building and shipping LLM-based agentic systems in production — real users, real incidents, real fixes.
- Several substantial applications built with Claude or comparable code models , with the context-engineering and verification habits described above.
- Hands-on LangGraph or an equivalent orchestration framework, and Amazon Bedrock or comparable LLM infrastructure; retrieval-augmented generation built for real use; and a clear grasp of agentic failure modes — hallucination, tool misuse, context loss.
- Strong Python , including async patterns and testing strategies for non-deterministic components.
- Experience leading a small team or workstream to dates that mattered , and taking a prototype through to production.
- Clear written and spoken English — you will be specifying behaviour for systems that behave probabilistically.
NICE TO HAVE
- Travel technology — booking, distribution, GDS or NDC — or another domain where a wrong answer has financial or regulatory consequence.
- Model Context Protocol or comparable agent-tool integration.
- AWS depth beyond Bedrock; exposure to Oracle Cloud.
- Responsible-AI or model governance practice.
- Voice or real-time conversational systems in production.
- Retrieval at scale — hybrid search, reranking, freshness under constant ingest.
PROBABLY NOT A FIT IF
- Your LLM work has been prompt design and integration rather than building systems.
- You are looking to step back from coding and lead through others — this pod is small and you are in it.
- You want a research role. This is production engineering with a research-shaped surface.
📌 Senior Product Engineer (Hyderabad)
🏢 Mondee
📍 Hyderabad