02 Aug
|
Zenalyst
|
India
Everyone is demoing agents. Very few are running them in production multi-tenant, cost-controlled, evaluated, auditable, and trusted with an enterprise's real workflows.
That gap is the job. And it sits at the centre of a multi-billion dollar industry that is reimagining how enterprises operate.
You can read about it. Or you can build the infrastructure it runs on.
Role : Lead AI / MLOps Engineer
Company : Zenalyst (zenalyst.ai)
Location : Bengaluru - 5 days from office
Experience : 4 - 6 years
Reports to : Engineering Manager
Education : BE / ME preferred
About Zenalyst :
Zenalyst is an enterprise workflow automation company. We build AI-native platforms that take the manual, fragmented, spreadsheet-and-email workflows running large organisations across finance, procurement, legal, operations and beyond and turn them into automated, intelligent, auditable systems.
Finance is one of the domains we have gone deep in. It is not the boundary. Wherever an enterprise process is slow, manual and high-stakes, that is our problem space.
We are funded, rapidly expanding, and already live with paying enterprise clients. Our agents are not a demo they run inside customer workflows every day.
We work with frontier models and with pretrained open-weight models from Hugging Face, and we choose deliberately between them on capability, cost, latency and data-control grounds.
The Role :
You will own the AI platform the layer everything intelligent in our product runs on.
That means model strategy and serving, agentic orchestration, retrieval, evaluation, guardrails, observability and cost. You will decide what runs on a frontier API and what runs on a fine-tuned open-weight model on our own infrastructure, and you will be accountable for the quality, latency, spend and safety of both.
This is a hands-on lead role. You will write code, set the standards, review the designs, and own delivery for a small high-calibre team.
What You Will Own :
Models & Serving :
- Own our model strategy across frontier APIs (Claude, GPT, Gemini, Bedrock, Azure OpenAI) and pretrained open-weight models from Hugging Face (Llama, Mistral, Qwen and similar)
- Fine-tune, adapt (LoRA / PEFT), quantise and deploy open-weight models where they beat an API call on cost, latency or data control
- Build and run inference infrastructure vLLM / TGI / Triton, GPU provisioning, autoscaling, batching, caching
- Model routing and tiering: the right model for the right task, with fallbacks that keep workflows alive
Agentic & Retrieval Systems :
- Productionise agentic workflows tool and function calling, multi-agent orchestration, planning loops, human-in-the-loop checkpoints, MCP-style tool integration
- Own the RAG stack end to end : ingestion, chunking, embeddings, vector search (pgvector and/or a dedicated vector store), hybrid retrieval, re-ranking, context assembly
- Design agents that execute enterprise workflows reliably retries, idempotency, state, failure containment
Evaluation & Quality :
- Build the evaluation harness : golden datasets, regression suites, LLM-as-judge, groundedness and hallucination checks
- Make model and prompt changes measurable no shipping on vibes
- Monitor drift, quality regressions and edge-case failures in production
Guardrails, Security & Compliance :
- PII detection and redaction, prompt injection defence, output validation, jailbreak resistance
- Tenant isolation for AI workloads, data residency, retention controls, and full audit trails of what an agent did and why
- Partner with engineering on enterprise security posture
Platform, Cost & Observability :
- MLOps foundations : model registry, versioning, reproducible pipelines, CI/CD for models and prompts
- Full AI observability traces, token telemetry, p95 latency, cost per workflow and per tenant
- Drive unit economics down as usage scales up
- Multi-cloud deployment (AWS / Azure) with Docker and Kubernetes, in partnership with DevOps
Delivery & Team :
- Work directly with the Product Owner for requirement clarity, prioritisation and task allocation
- Own sprint delivery: scope, sequence, unblock, ship on commitment
- Mentor engineers and raise the technical bar around you
- Prioritise ruthlessly know what ships now, what waits, and why
Our Environment :
- Core product : Java, Spring Boot, Hibernate, React
- AI / ML : Python, PyTorch, Hugging Face, LangChain / LlamaIndex or equivalent,
vLLM
- Data : PostgreSQL ( pgvector), MongoDB, S3, Kafka
- Cloud : AWS, Azure (Bedrock / SageMaker / Azure ML)
- DevOps : Docker, Kubernetes, Git, CI/CD
What We Are Looking For :
Must have :
- 4 -6 years in engineering, with substantial recent time on ML / AI systems in production
- Strong Python; comfortable integrating with a Java/Spring Boot service landscape
- Hands-on with frontier model APIs AND with pretrained Hugging Face models fine-tuning, serving and evaluating both
- Production experience with LLM applications : RAG, agents, tool calling, prompt and context engineering
- Real MLOps depth : deployment pipelines, versioning, monitoring, reproducibility
- Inference optimisation quantisation, batching, caching, GPU efficiency, latency and cost tuning
- Evaluation discipline you can prove a model change made things better
- Multi-cloud (AWS and/or Azure), Docker, Kubernetes
- Working knowledge of data, ETL, pipelines and streaming (Kafka)
- Genuine ownership you chase problems to closure without being asked
Good to have :
- Built an AI platform zero to one inside an enterprise product
- Multi-tenant SaaS experience and AI security / guardrail design
- Financial services, banking or other regulated-domain exposure
- Integration experience with SAP, ERP or CRM systems
- Familiarity with SOC 2 / ISO 27001 style controls for AI systems
- Open-source contributions, papers, or published model work
Education :
- BE / ME (or equivalent engineering degree) preferred
- Candidates from premium institutes IIT, IIIT, NIT, BITS are a strong plus
Why This Role :
- You own the AI platform. Not one model, not one feature the layer the whole product depends on.
- Frontier and open-weight, both. Real freedom to choose the right tool, and real accountability for the choice.
- Production, not prototypes. Your agents run inside enterprise workflows with money and compliance attached.
- Funded, with live clients. Validation and stability, with greenfield scope.
- Clear ladder. We are expanding fast; the engineers who build the platform will lead the org that follows.
- Direct access. Small team, short decision paths, leadership one conversation away.
Ready to Build?
If you have taken AI from notebook to production and kept it fast, cheap, secure and measurably good we should talk.
📌 Zenalyst.AI - Lead AI/MLOps Engineer (India)
🏢 Zenalyst
📍 India