24 Sep
|
Impetus Technologies
|
Bengaluru
24 Sep
Impetus Technologies
Bengaluru
We are building a next-generation AI Observability and Trust Governance platform that helps enterprises monitor, evaluate, and govern their GenAI applications. We are looking for an experienced Architect (8-12 years) who can blend architectural depth with a product development mindset to drive the design and innovation roadmap for new platform capabilities.
This is a hands-on architecture role rather than a pure strategy or pure-development position. You will spend your time on design thinking, capability evaluation, prototyping, and guiding engineering execution - not on being the primary hands-on coder for the team. You will also bring an understanding of DevOps practices and cloud/data platforms, since observability and trust capabilities are deeply intertwined with CI/CD, infrastructure, data pipelines, and operational processes.
Above all, we are looking for someone with genuine curiosity and passion for the GenAI space - someone who is always exploring new models, tools, and techniques, and enjoys staying ahead of a quick-moving field.
Roles & Responsibilities
Architecture & Technical Leadership
- Own the end-to-end architecture for the AI observability platform - covering tracing, logging, evaluation, token consumption, monitoring, and governance layers for LLM/GenAI and traditional ML workloads.
- Design scalable, extensible frameworks for capturing model telemetry (prompts, completions, latency, token usage, cost, embeddings, tool calls, agent traces) across multi-model and multi-agent systems.
- Define reference architectures for trust and safety capabilities: hallucination detection, bias/fairness monitoring, PII/data leakage detection, prompt-injection defense, guardrails, and drift detection.
- Establish architectural standards, design patterns, and best practices for the engineering team to follow, and review designs proposed by engineers for soundness and scalability.
Innovation & Capability Development
- Continuously scan and evaluate the fast-evolving GenAI observability landscape (LLM eval frameworks,
OpenTelemetry GenAI semantic conventions, guardrail libraries, model risk/governance tools) and make build-vs-buy-vs-integrate recommendations.
- Drive new capabilities from concept through proof-of-concept to production readiness - e.g., RAG evaluation, agentic workflow tracing, automated red-teaming, model cards, AI bill-of-materials / lineage tracking.
- Track regulatory and industry trends (EU AI Act, NIST AI RMF, ISO 42001) and translate them into practical product features and controls.
- Represent the platform in technical forums, evaluate emerging open-source and commercial tools, and contribute to internal IP and thought-leadership content.
Product & Cross-Functional Collaboration
- Partner closely with Product Management to shape the roadmap, write technical specs/RFCs, and prioritize capabilities based on customer and market needs.
- Mentor engineers, lead design reviews, and help raise the overall technical and architectural bar across the team.
Guided, Hands-on Involvement (not primary development)
- Prototype and validate ideas quickly in Python to prove out architectural concepts before handing off to the engineering team for full-scale build.
- Review critical SDK/API designs and integration frameworks for developer experience, ensuring customers can instrument their AI applications with minimal friction.
- Get hands-on selectively - spikes, POCs, and complex debugging - rather than owning day-to-day feature development.
Key Technical Skills GenAI & ML Expertise
- Strong conceptual and applied understanding of LLM architectures, RAG pipelines, agentic systems (multi-agent orchestration, tool/function calling),
fine-tuning, and prompt engineering.
- Familiarity with LLM evaluation methodologies - LLM-as-judge, embedding-based similarity, hallucination/faithfulness scoring, red-teaming, and adversarial testing.
- Good understanding of token optimization
- Working knowledge of responsible AI concepts: bias/fairness metrics and AI governance frameworks (NIST AI RMF, EU AI Act, ISO 42001).
- Exposure to vector databases, embeddings, and semantic search concepts.
Observability, Platform Engineering & DevOps
- Practical experience with observability tooling and standards - OpenTelemetry (including GenAI semantic conventions), distributed tracing, logging/metrics pipelines.
- Awareness of LLM observability tools such as LangSmith, Langfuse, Arize Phoenix, TruLens, Traceloop/OpenLLMetry, Datadog LLM Observability
- Exposure to guardrail/safety libraries (Guardrails AI, NeMo Guardrails, Llama Guard, Presidio for PII detection).
- Understanding of distributed systems concepts - event-driven architectures, streaming (Kafka), scalable storage (time-series DBs, columnar stores).
- Exposure to data engineering/platform tools (e.g., Databricks, Spark) for processing large volumes of telemetry, evaluation, and model performance data at scale.
Programming & Engineering
- Good working proficiency in Python - enough to prototype confidently, review code meaningfully, and communicate credibly with engineers; deep day-to-day production coding is not the primary expectation of this role.
- Familiarity with GenAI SDKs/frameworks such as LangChain, LlamaIndex, and OpenAI/Anthropic/Azure SDKs.
- Hands-on experience working in cloud environments (AWS/Azure/GCP)
- Exposure to modern data platforms such as Databricks (or equivalent - Snowflake, EMR) for large-scale data processing, feature/telemetry pipelines, and analytics that feed observability and evaluation workloads.
- Working knowledge of modern microservices/API design principles (REST/gRPC).
📌 AI Observability & Trust Architect (Bengaluru)
🏢 Impetus Technologies
📍 Bengaluru