06 Oct
|
HCL Comnet
|
Noida
- Experience: 12-15years
- Location: Noida/Bangalore
- Requirements: Experience in Observability and OpenTelemetry - LLM/Applications
- Send resumes to: [HIDDEN TEXT] with below details:
- Name:
- Exp:
- CTC:
- ECTC:
- Notice period:
- Current location:
Job Description:
Level: 12+ years in software or platform engineering, including 3+ years designing
GenAI systems that other engineers shipped.
Own the architecture for LLM applications and the platform under them: models,
data, agents, evaluations, delivery, and the controls around them.
You will
Set the architecture for model access, retrieval, agent runtime, shared tools,
and one place to see traces and scores.
Choose when one agent is enough and when work splits across agents. MCP
is the tool boundary. Multi-agent handoff, including A2A, is the agent
boundary.
Define RAG so index lifecycle, citations, and freshness are part of the design,
and judge whether the enterprise data can support it.
Set standards others implement: OpenTelemetry GenAI as the span contract,
an evaluation approach so scores stay comparable, and LLMOps so prompts,
agents, and indexes are versioned, regression-tested, and rollback-able.
Place guardrails, tenancy, secrets, and retention in the platform. Decide what
an agent may read or change, and where a person must approve.
Shape inference for latency and cost: model routing, smaller models where
they are enough, caching,
and batching where they help.
Write the architecture down and defend the tradeoffs with engineering and
with the people who own the business outcome.
Skills
A system you can walk through from request to stored trace, score, and cost,
on a cloud you have operated: AWS, Azure, or GCP, including Bedrock, Azure
OpenAI, Vertex AI, or Databricks.
One agent framework at design depth: LangGraph, CrewAI, AutoGen,
Semantic Kernel, OpenAI Agents SDK, or Google ADK. Familiarity with the
others is enough.
RAG architecture: hybrid retrieval, reranking, grounding, and the data work
underneath the index.
MCP and multi-agent design, including memory, tool scope, and human
approval. Direct A2A experience is a plus.
Evaluation and observability standards: golden sets, live signals, LLM-as-
judge, and GenAI spans for model, retrieval, tool, and agent steps. Operating
experience with MLflow, Arize Phoenix, LangSmith, or Langfuse is the
evidence. You will set the gen_ai.* contract here. You do not need to have
authored that spec elsewhere.
LLMOps: CI/CD for prompts and agents, regression gates, model and index
rollout, and rollback.
Platform controls: isolation between tenants and agents, guardrails, audit,
retention, and token cost.
Safety and reliability: prompt injection, tool abuse, timeouts, and a defined
path when the model is wrong.
📌 Gen AI Architect (Noida)
🏢 HCL Comnet
📍 Noida