03 Sep
|
Leena AI
|
Gurugram
About Leena AI
Leena AI is a leader in Agentic AI for the enterprise. We are building an iconic company, delivering AI Colleagues that transform back-office functions and accelerate the full promise of Generative AI - unlocking real productivity gains, cutting costs, and delighting employees at scale.
Leena AI provides the most forward-looking, open, and scalable Agentic AI architecture for the enterprise - it empowers CIOs and CTOs to develop, deploy, and manage AI Colleagues for the back office at scale. Built with full governance, compliance, security, and auditability at its core.
Leena AI integrates with 1000 applications, including SAP, Salesforce, ServiceNow, Workday, and Microsoft Office 365. We are proud to be trusted by 500 global enterprises and 20 million employees, including leading brands such as Nestlé, Puma, Coca-Cola, Sony, and Etihad Airways.
Founded in 2018 and headquartered in New York, Leena AI has secured over $40M in financing from top-tier investors including Greycroft, Bessemer Venture Partners, B Capital, and Y Combinator.
Role Overview
Leena AI’s agentic assistant answers employees’ HR, IT, and policy questions from each customer’s knowledge base, across ~400 enterprise tenants.
This role leads the Knowledge Intelligence team: retrieval and ranking, answer grounding, conflict detection, LLM policy, and the evaluation infrastructure that keeps an LLM-powered agent trustworthy at enterprise scale.
The role blends people leadership, system design, and hands-on ownership of production LLM systems. The candidate we hire has shipped LLM-based products to real users — with evals, tool calling, and cost/latency engineering not only classical ML. Classical ML depth is necessary here; it is not sufficient.
LLM & Agentic Systems Ownership
- Own production LLM systems end to end: RAG pipelines, agent tool-calling flows, model selection and routing, guardrails, and answer-vs-escalate decision logic.
- Build and own evaluation infrastructure: golden datasets per tenant, LLM-as-judge graders calibrated against human judgment, regression gates on every prompt and model change, and online quality monitoring.
- Drive answer quality: grounding and attribution against ACL-visible sources, hallucination control, and confidence calibration across heterogeneous tenant knowledge bases.
- Engineer cost and latency at scale: routing to right-sized models, caching,
and token budgets — with per-query economics you can state from memory.
Retrieval & ML Ownership
- Own search quality across tenants: embeddings, rerankers, hybrid retrieval, chunking/sectioning, and measurement (recall@k, MRR) — including per-tenant diagnosis when quality regresses.
- Guide the full model lifecycle: data preparation, training, evaluation, deployment, monitoring, and iteration — including the data-bias traps in learning from your own system’s logs.
- Ensure ML systems are reliable, observable, scalable, and cost-productive in production.
Engineering Leadership
- Lead, mentor, and grow a cross-functional team of backend and ML engineers.
- Create a high-trust, high-ownership culture focused on quality, learning, and delivery.
- Support career development, performance management, and technical growth of team members.
Architecture & Technical Direction
- Drive architecture for distributed, multi-tenant backend systems that integrate LLM inference, retrieval, and data pipelines.
- Separate what must be learned from what should be a rule: design systems where models handle judgment and policy handles constraints.
- Partner with senior ICs on trade-offs across accuracy, latency, cost, and system complexity; review designs and code across services, ML pipelines, and infrastructure.
Execution & Cross-Functional Collaboration
- Plan and execute multi-quarter roadmaps spanning product engineering and ML initiatives; balance experimentation velocity with enterprise-grade reliability and compliance.
- Work closely with Product, Applied AI, and Customer teams to align technical execution with business outcomes; translate product requirements into feasible plans.
- Represent engineering in roadmap, prioritization, and stakeholder discussions — with metrics, not adjectives.
RequirementsLLM Production Experience (hard requirement)
- Shipped LLM-based systems to production with real users — RAG, agents, or LLM-powered product features — and operated them: this is the bar, not a bonus.
- Designed and run evaluation for LLM systems: offline golden sets, judge models validated against human labels, and eval-gated rollout of prompt or model changes.
- Debugged real agent and tool-calling failures (malformed calls, loops, hallucinated grounding) and can tell those stories concretely, with numbers.
- Reasoned about LLM system economics in production: cost per query, latency budgets, routing and caching strategies.
Experience & Background
- 7 years of software engineering experience, with prior hands-on IC work in ML or applied AI systems.
- 2 years managing engineers, ideally including ML or data-focused teams.
- Experience building and operating production ML systems (not just experimentation or research).
Technical Expertise
- Strong retrieval/NLP foundation: embeddings, rerankers, hybrid search, and search-quality measurement.
- Strong foundation in distributed systems and backend engineering (Node.js or Python, APIs, cloud infrastructure); multi-tenant SaaS experience valued.
- Ability to reason about ML-specific trade-offs: accuracy vs latency, offline vs online metrics, calibration, retraining strategies.
Leadership & Judgment
- First-principles problem framing: interrogates objectives, cost asymmetries, and constraints before proposing architecture. Our interviews test this habit directly.
- Proven ability to lead teams through ambiguity and technically complex problem spaces.
- Comfortable being hands-on when needed, while primarily operating as a multiplier for the team.
- Excellent communication with technical and non-technical stakeholders; can explain LLM behavior and risk to product and business partners.
How We Interview
We tell candidates this up front because it produces better conversations. Expect:
(1) a deep-dive on an LLM system you personally shipped — architecture, eval story, failures, and cost;
(2) a design problem from our domain, where asking the right questions is scored.
Education
Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field (Tier 1 preferred).
Nice to Have
- Conversational AI, enterprise knowledge management, or search at multi-tenant scale.
- Classical recommender/personalization systems experience (valued in addition to, not instead of, LLM production work).
- Prior experience scaling AI systems in B2B SaaS or regulated, high-availability environments.
📌 Engineering Manager - AL/ML (Gurugram)
🏢 Leena AI
📍 Gurugram