Senior AI Engineer – RAG & LLM Systems (Bengaluru)

Senior AI Engineer – RAG & LLM Systems (Bengaluru)

12 Aug
|
FICO
|
Bengaluru

12 Aug

FICO

Bengaluru

AI Engineer, Incident Intelligence Platform We are looking for an AI Engineer with strong hands-on skills in LLM systems and retrieval-augmented generation to join our incident intelligence team. In this role, you will architect and operate the full pipeline that powers our AI-driven root cause analysis (RCA) platform, from the vector database and retrieval layer through to Claude prompt engineering, structured output validation, and confidence scoring and intelligent routing. You will build the systems that decide whether an RCA is accurate enough to auto-notify an on-call engineer or needs human review, working across our multi-tenant PaaS platform hosting 100+ enterprise customer workloads on shared AWS EKS.

Qualifications:

4+ years of experience in AI/ML engineering, data engineering, NLP, or a closely related field, including hands-on LLM work.

Production experience building and operating a vector database (pgvector, Pinecone, Weaviate, or equivalent).

Strong understanding of embedding models and how embedding quality affects retrieval accuracy.

Experience designing and operating data ingestion pipelines at scale (Celery, Airflow, or similar).

Demonstrated experience prompt engineering for production LLM systems requiring reliable structured outputs, not just chatbots or demos.

Experience building LLM evaluation frameworks, including golden test sets, judge prompts, and accuracy metrics.

Robust Python skills, including FastAPI, Pydantic, SQLAlchemy, and async patterns.

Clear understanding of how RAG context affects LLM reasoning and how to diagnose retrieval-driven errors.

Familiarity with semantic search concepts: cosine similarity, approximate nearest neighbor, and hybrid search.

Production experience with the Claude API (messages format, system prompts, structured output patterns) is a plus.





Experience with LLM output validation frameworks such as Pydantic or Guardrails is a plus.

Understanding of confidence calibration techniques (ECE, Platt scaling, temperature effects) is a plus.

Exposure to SRE or observability domains, and familiarity with multi-tenant architectures where tenant isolation is a hard requirement, is a plus.

Knowledge of re-ranking models (cross-encoders) and experience with Kubernetes/AWS EKS. Key Responsibilities:

Vector Database & Retrieval: Design and operate the pgvector schema across all collections (past incidents per tenant, runbooks, operational context, CI/CD change events), owning indexing strategy, query performance tuning, and strict data isolation across 100+ tenants with zero cross-tenant leakage.

Embedding & Ingestion Pipelines: Build and maintain embedding pipelines that convert operational documents into searchable vector representations, selecting embedding models and designing chunking strategies for structured JSON, unstructured prose, and time-series data types.

Semantic Search: Build the RAG retrieval layer that runs at incident time, tuning relevance thresholds, re-ranking strategies, and hybrid (dense + sparse) search to surface the most contextually relevant documents per collection.

Ingestion Architecture: Design the full ingestion lifecycle, including runbook imports, historical incident backfill, CI/CD post-deploy event hooks,



and the feedback loop that embeds resolved incidents back into the vector database.

Claude Prompt Engineering: Design and iterate the prompt that assembles Grafana's stateless RCA, live signal data, and RAG-retrieved context into a single input, instructing Claude to reason across evidence types and produce a structured, actionable RCA readable in under 30 seconds.

Structured Output Validation: Define and maintain the JSON schema for RCA output (root cause, confidence score, contributing factors, immediate actions, SLA impact, precedent reference, escalation path) and build the Pydantic validation layer that catches schema violations before delivery.

Confidence Scoring: Own the calibrated 0–1 confidence scoring framework weighting RAG precedent strength, signal consistency, deployment correlation, runbook recognition, root cause specificity, and customer context fit, and build the FastAPI gate logic that routes RCAs to the correct action tier.

QA Validation Loop: Maintain a golden test set of 20–30 historical incidents and run the full pipeline against it after every prompt or RAG change, using a judge-prompt approach to score accuracy and catch regressions.

Confidence Calibration: Run monthly calibration checks comparing confidence scores against SRE-confirmed outcomes, calculating Expected Calibration Error (ECE) and adjusting thresholds accordingly.

Versioning & Cost Monitoring: Build a prompt versioning system with changelog tracking, and monitor token usage, cost per RCA, and API latency, alerting on unexpected cost trends.

Data Quality Metrics: Track retrieval quality metrics including relevance scores, coverage gaps, cross-tenant isolation verification, and ablation testing of RCA quality with and without each data collection.

📌 Senior AI Engineer – RAG & LLM Systems (Bengaluru)
🏢 FICO
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior ai engineer – rag & llm systems (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: senior ai engineer – rag & llm systems (bengaluru) / bengaluru