08 Oct
|
Brainvire Infotech
|
Mumbai
08 Oct
Brainvire Infotech
Mumbai
AI ENGINEER
Location: Mumbai
Experience: 2 to 5 years, with at least 2 years building LLM systems that
ran in production
ABOUT THE ROLE
You will own an AI service from the retrieval code through to the token
bill it generates. We build agentic and RAG systems for clients in
e-commerce, healthcare, financial services and enterprise SaaS. These run
under real load with real money attached, so the job splits roughly into
building the system, proving it behaves correctly, and keeping it fast
and affordable.
This is a senior individual contributor role. You will review other
people's code and carry associate engineers.
WHAT YOU WILL DO/OWN
- Build multi-agent and RAG systems on LangGraph, with agent memory,
tool permissioning, spend caps and human approval on anything
irreversible.
- Design the retrieval layer and be able to say what its recall and
precision are, separately from answer quality.
- Own evaluation. Build golden sets from production traffic, run them in
CI, and decide what score change blocks a release.
- Instrument tracing and monitoring so that when a client says the
answer was wrong, you can find the request and explain why.
- Run the serving layer. Latency budgets, caching, batching, model
routing, and a cost per request you can quote without looking it up.
- Defend the system against prompt injection, data leakage across
tenants, and jailbreaks, and test those defences rather than assume them.
- Integrate with client systems: Salesforce, SAP, ServiceNow, Magento,
Odoo, and whatever else the account runs on.
- Work directly with client engineering teams during delivery.
WHAT YOU NEED
Engineering: Advanced Python with async and typing, Production API design
in FastAPI or Django REST, SQL and data modelling, Docker,
Git and code review, Linux debugging, testing systems whose output
changes between runs.
Models: OpenAI, Anthropic and Gemini APIs hands on. Amazon Bedrock, Azure
AI Foundry or Google Vertex AI. STT/TTS. Vision and image models including CLIP, SAM,
Stable Diffusion or FLUX. Tool calling with schema enforcement,
structured output, streaming, and context and token budget management. Context Engineering
Agents & orchestration: LangGraph and LangChain, plus one of CrewAI, AutoGen, OpenAI
Agents SDK or Pydantic AI. MCP. Idempotency and compensating actions for tools that change state.
Retrieval & knowledge: RAG, chunking strategy, embedding selection, hybrid
search, reranking, permission-aware retrieval, document parsing.
PostgreSQL with pgvector and one dedicated vector database at scale
(Qdrant, Weaviate, Milvus or Pinecone), GraphRAG with Neo4j.
Evaluation and safety: Ragas or DeepEval, LLM as judge with calibration,
regression gates in CI, LangSmith or Langfuse, OpenTelemetry, drift
monitoring, PII redaction with Presidio, prompt injection defence.
Serving & performance: Latency budgeting — p50 / p95 / p99, Semantic caching, batching, model cascade and routing, and cost instrumentation.
Machine learning: scikit-learn, XGBoost, Feature engineering, Time series — ARIMA, Prophet, LSTM, Hugging Face Transformers, PEFT with LoRA or QLoRA, and judgement to know when a classical model beats an LLM.
Good to have: WebRTC and LiveKit for voice, vLLM / TGI / Triton, NeMo Guardrails / Llama Guard, Open-weight models such as Llama, Mistral or Qwen, PyTorch, OpenCV, Semantic Kernel
Pay: Up to ₹1,000,000.00 per year
Perks:
- Flexible schedule
- Paid sick time
- Provident Fund
Work Location: In person
📌 AI Engineer (Mumbai)
🏢 Brainvire Infotech
📍 Mumbai