AI Observability & Trust Architect (India)

AI Observability & Trust Architect (India)

23 Sep
|
Impetus Technologies
|
India

23 Sep

Impetus Technologies

India

Qualification:

We are building a next-generation AI Observability and Trust Governance platform that helps enterprises monitor, evaluate, and govern their GenAI applications. We are looking for an experienced Architect (8-12 years) who can blend architectural depth with a product development mindset to drive the design and innovation roadmap for new platform capabilities.

This is a hands-on architecture role rather than a pure strategy or pure-development position. You will spend your time on design thinking, capability evaluation, prototyping, and guiding engineering execution - not on being the primary hands-on coder for the team. You will also bring an understanding of DevOps practices and cloud/data platforms, since observability and trust capabilities are deeply intertwined with CI/CD, infrastructure, data pipelines, and operational processes.

Above all, we are looking for someone with genuine curiosity and passion for the GenAI space - someone who is always exploring new models, tools, and techniques, and enjoys staying ahead of a fast-moving field.

Skills Required:

DevOps,GenAI,Observability,Cloud

Role:

Architecture & Technical Leadership

- Own the end-to-end architecture for the AI observability platform - covering tracing, logging, evaluation, token consumption, monitoring, and governance layers for LLM/GenAI and traditional ML workloads.
- Design scalable, extensible frameworks for capturing model telemetry (prompts, completions, latency, token usage, cost, embeddings, tool calls, agent traces) across multi-model and multi-agent systems.
- Define reference architectures for trust and safety capabilities: hallucination detection, bias/fairness monitoring, PII/data leakage detection, prompt-injection defense, guardrails, and drift detection.
- Establish architectural standards, design patterns, and best practices for the engineering team to follow, and review designs proposed by engineers for soundness and scalability.

Innovation & Capability Development

- Continuously scan and evaluate the fast-evolving GenAI observability landscape (LLM eval frameworks, OpenTelemetry GenAI semantic conventions, guardrail libraries, model risk/governance tools) and make build-vs-buy-vs-integrate recommendations.
- Drive current capabilities from concept through proof-of-concept to production readiness - e.g., RAG evaluation, agentic workflow tracing, automated red-teaming, model cards, AI bill-of-materials / lineage tracking.
- Track regulatory and industry trends (EU AI Act, NIST AI RMF, ISO 42001)



and translate them into practical product features and controls.
- Represent the platform in technical forums, evaluate emerging open-source and commercial tools, and contribute to internal IP and thought-leadership content.

Product & Cross-Functional Collaboration

- Partner closely with Product Management to shape the roadmap, write technical specs/RFCs, and prioritize capabilities based on customer and market needs.
- Mentor engineers, lead design reviews, and help raise the overall technical and architectural bar across the team.

Guided, Hands-on Involvement (not primary development)

- Prototype and validate ideas quickly in Python to prove out architectural concepts before handing off to the engineering team for full-scale build.
- Review critical SDK/API designs and integration frameworks for developer experience, ensuring customers can instrument their AI applications with minimal friction.
- Get hands-on selectively - spikes, POCs, and complex debugging - rather than owning day-to-day feature development.

Key Technical Skills

GenAI & ML Expertise

- Strong conceptual and applied understanding of LLM architectures, RAG pipelines, agentic systems (multi-agent orchestration, tool/function calling), fine-tuning, and prompt engineering.
- Familiarity with LLM evaluation methodologies - LLM-as-judge, embedding-based similarity, hallucination/faithfulness scoring, red-teaming, and adversarial testing.
- Good understanding of token optimization
- Working knowledge of responsible AI concepts: bias/fairness metrics and AI governance frameworks (NIST AI RMF, EU AI Act, ISO 42001).
- Exposure to vector databases, embeddings, and semantic search concepts.

Observability, Platform Engineering & DevOps

- Practical experience with observability tooling and standards - OpenTelemetry (including GenAI semantic conventions), distributed tracing, logging/metrics pipelines.
- Awareness of LLM observability tools such as LangSmith, Langfuse, Arize Phoenix, TruLens, Traceloop/OpenLLMetry, Datadog LLM Observability
- Exposure to guardrail/safety libraries (Guardrails AI, NeMo Guardrails, Llama Guard, Presidio for PII detection).




- Understanding of distributed systems concepts - event-driven architectures, streaming (Kafka), scalable storage (time-series DBs, columnar stores).
- Exposure to data engineering/platform tools (e.g., Databricks, Spark) for processing large volumes of telemetry, evaluation, and model performance data at scale.

Programming & Engineering

- Good working proficiency in Python - enough to prototype confidently, review code meaningfully, and communicate credibly with engineers; deep day-to-day production coding is not the primary expectation of this role.
- Familiarity with GenAI SDKs/frameworks such as LangChain, LlamaIndex, and OpenAI/Anthropic/Azure SDKs.
- Hands-on experience working in cloud environments (AWS/Azure/GCP)
- Exposure to modern data platforms such as Databricks (or equivalent - Snowflake, EMR) for large-scale data processing, feature/telemetry pipelines, and analytics that feed observability and evaluation workloads.
- Working knowledge of modern microservices/API design principles (REST/gRPC).

Soft Skills & Mindset

- Product development mindset - comfortable with ambiguity, able to translate customer and market problems into clear technical requirements and phased roadmaps.
- Strong systems and design thinking - able to zoom out to architecture while staying grounded in practical implementation trade-offs.
- Curiosity and a bias for continuous learning - genuinely enjoys tracking a fast-moving field and separating hype from substance.
- Excellent communication skills
- Collaborative influence - able to drive alignment and technical decisions across Product, Engineering, DevOps, and Security without relying on positional authority.
- Mentorship orientation - invested in growing engineers' capabilities through design reviews, pairing, and constructive feedback.
- Pragmatic decision-making - balances innovation with delivery timelines and knows when to prototype versus when to commit to a full build.
- Ownership and accountability - takes end-to-end responsibility for architectural outcomes, even across team boundaries.
- Adaptability - stays effective as priorities shift in a fast-evolving, less-defined problem space.
- Genuine passion for exploring new technologies - proactively tries out new GenAI models, tools, and techniques as they emerge, rather than waiting to be asked.
- A perpetual-learner attitude - actively follows GenAI research, product launches, and industry developments, and brings relevant findings back to the team and roadmap.

Experience:

8 to 10 years

Job Reference Number:

14099

📌 AI Observability & Trust Architect (India)
🏢 Impetus Technologies
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: ai observability & trust architect (india) / india