Senior AI Evaluation & Reliability Engineer (Ahmedabad)

Senior AI Evaluation & Reliability Engineer (Ahmedabad)

03 Sep
|
Aubergine Solutions
|
Ahmedabad

03 Sep

Aubergine Solutions

Ahmedabad

Role Overview

We are looking for a Senior AI Evaluation Reliability Engineer who is passionate about solving one of the most important challenges in AI: How do we know an AI system is actually working, improving, and delivering business value. You will design and build production-grade evaluation systems for LLMs, RAG applications, and multi-agent systems, while helping enterprise clients and our engineering teams adopt AI with confidence.

This role goes beyond building evals. You will be a trusted AI consultant to clients, a technical mentor to engineers, and a key contributor to our journey towards becoming an AI superagency.

What You Will Own

- Make AI Performance ROI Measurable: Build evaluation strategies that connect AI performance to business outcomes and ROI.

- Define quality benchmarks, SLOs, risk thresholds, and success criteria.

- Track hallucinations, reliability issues, and production risks.

- Build executive-friendly AI quality and ROI scorecards.

- Help clients determine where AI should be autonomous, supervised, or avoided.

- Don't just measure model accuracy. Measure business impact.

Build Production-Grade Evaluation Pipelines

- Architect automated evaluation pipelines for LLMs, RAG, and agentic systems.

- Integrate evaluations into CI/CD using GitHub Actions, GitLab CI, or equivalent.

- Build regression suites for prompts, models, tools, and workflows.

- Measure faithfulness, context precision, answer relevance, semantic drift, and task completion.

- Establish statistically meaningful benchmarks and continuously monitor AI quality.

- Evaluation should become part of engineering, not a final QA step.

Engineer LLM-as-a-Judge Systems

- Design reference-based and reference-free LLM evaluation frameworks.





- Create structured rubrics and scoring systems.

- Build calibration loops using human-labelled datasets.

- Identify and mitigate judge biases such as position, verbosity, and self-preference bias.

- Measure judge reliability and optimize evaluation quality, latency, and cost.

Own Evaluation Economics

- LLM evaluations can become expensive quickly.

- You will: Design tiered evaluation strategies.

- Use deterministic and heuristic graders for simple checks.

- Reserve powerful LLM judges for complex evaluations.

- Optimise batching, concurrency, and parallel execution.

- Track evaluation cost and balance quality, latency, and inference spend.

Benchmark RAG Agentic Systems

- Build systematic benchmarks across:

- Retrieval quality, faithfulness, and answer relevance

- Embedding, chunking and re-ranking strategies

- Vector databases and retrieval architectures

- Agent task completion and tool-calling accuracy

- State retention, multi-step workflows and failure recovery

- Latency, reliability and cost

- Build synthetic datasets, golden test suites, and adversarial scenarios to continuously expand evaluation coverage.

Be a Trusted AI Consultant

- You will work directly with North American enterprise clients as a technical AI advisor.

- You will: Lead AI architecture and evaluation discussions.





- Translate complex AI metrics into clear business recommendations.

- Define AI quality, reliability, and risk frameworks.

- Present evaluation telemetry and ROI scorecards to technical and business stakeholders.

- Challenge assumptions and recommend the right AI solution, even when that means saying "don't use AI here."

- You should be equally comfortable discussing LLM evaluation with engineers and ROI with a CTO/COO/CEO/CFO.

Drive AI Consulting Presales

Your consulting mindset will extend beyond delivery. You will actively participate in presales and new business opportunities, helping us win strategic AI engagements with enterprise clients.

- Participate in discovery calls, solution workshops, and technical discussions with prospective clients.

- Understand client challenges and identify where AI, agents, RAG, or automation can create meaningful business value.

- Shape AI solution approaches, evaluation strategies, and technical proposals.

- Contribute to statements of work, estimates, solution architectures, and project proposals.

- Build compelling technical narratives and demonstrations that showcase our AI capabilities.

- Present AI solutions and recommendations to technical and executive stakeholders.

- Help identify opportunities to expand existing engagements through recent AI capabilities.

- Partner closely with Sales, Delivery, Product, and Engineering teams to convert opportunities into successful engagements.

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Senior AI Evaluation & Reliability Engineer (Ahmedabad)
🏢 Aubergine Solutions
📍 Ahmedabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior ai evaluation & reliability engineer (ahmedabad) / ahmedabad

Subscribe to this job alert:

Get the latest job offers by email for: senior ai evaluation & reliability engineer (ahmedabad) / ahmedabad