Architect (Noida)

Architect (Noida)

24 Sep
|
HCLTech
|
Noida

24 Sep

HCLTech

Noida

We are hiring for AI Model Architect.

Location: Noida/ Chennai/ Bangalore/ Hyderabad/ Pune

Experience: 10+ years

Key Responsibilities

AI Architecture & SLM Development

- Own enterprise GenAI architecture and build domain SLMs using Gemma, Llama, Qwen, Phi, Mistral, or equivalent models.
- Deliver multi-turn operational workflows for triage, diagnosis, approval-gated remediation, validation, ticket updates, escalation, durable state, resumability, and loop prevention.

Data Preparation & Ingestion

- Build event-driven and scheduled ingestion for PDF, DOCX, PPTX, XLSX, SharePoint, APIs, and approved email folders or attachments using change notifications, delta sync, checkpoints, and reconciliation.
- Implement idempotency, incremental reprocessing, OCR/parsing, deduplication, malware and PII/secret controls, metadata and ACL preservation, security-trimmed retrieval, durable queues, retries, dead-letter handling, lineage, freshness, and ingestion monitoring.

Fine-Tuning & Alignment

- Lead SFT/instruction tuning, continued pretraining, domain adaptation, LoRA, QLoRA, DoRA, DPO/RLHF, knowledge distillation, quantization-aware optimization, and model compression.
- Engineer training datasets with cleaning, labeling, synthetic data, JSONL/chat formats, versioning, quality gates, and train/validation/test splits.

RAG & Grounded AI

- Architect RAG with chunking, embeddings, hybrid dense/sparse search, metadata filtering, reranking, context assembly, citations, and grounded-ness evaluation using Qdrant or equivalent.
- Build GraphRAG with Neo4j/AuraDB, Amazon Neptune, etc.; apply ontologies, entity resolution, graph traversal, subgraph retrieval, Cypher/open Cypher or Gremlin, and vector-graph hybrid retrieval for multi-hop reasoning and provenance.

Inference Engineering & Model Serving

- Deliver low-latency serving with vLLM, TensorRT-LLM, SGLang, TGI, or Triton using agile batching, paged attention, KV/prefix caching, speculative decoding,



streaming, routing, autoscaling, and INT4/INT8/FP8, AWQ, or GPTQ quantization.
- Develop intelligent model routing across SLMs, specialist models, and frontier LLMs using rules, semantic or complexity classification, model cascades, and learned routing; select the lowest-cost model that meets task-fit, quality, latency, context, privacy, safety, region, and availability requirements, with budgets, token metering, caching, circuit breakers, fallback, and continuous evaluation of routing accuracy, escalation rate, cost per successful task, and quality or latency regression.
- Declare and test TTFT, inter-token and p50/p95/p99 end-to-end latency, tokens/second, requests/second, concurrency, queue time, error rate, GPU/KV-cache utilization, context and token limits, cost, and performance under sustained, burst, failover, and degraded modes.

API, Platform & Production Operations

- Build OpenAI-compatible APIs with structured outputs, function/tool calling, orchestration, API gateway, authentication, RBAC/ACL, rate limits, quotas, secrets, and audit logging.
- Deploy through Docker, Kubernetes, Helm, Terraform, and CI/CD across Azure/AWS/GCP with multi-node/multi-zone HA, load balancing, autoscaling, automated failover, graceful degradation, blue-green/rolling release, rollback, cross-region DR, backups, point-in-time recovery, and tested RTO/RPO.

Evaluation, Security & Governance

- Define release gates for task success, Exact Match/F1, retrieval precision/recall, grounded-ness, citation coverage, hallucination/refusal, schema validity,



safety, latency, throughput, availability, RTO/RPO, and cost.
- Implement layered guardrails across input, retrieval, dialogue, routing, tool execution, and output: prompt-injection/jailbreak defense, PII/secret masking, grounding checks, tool allowlists, least privilege, HITL/dual approval, blast-radius and retry limits, kill switches, and deterministic fallback.
- Establish OpenTelemetry-compatible model/agent tracing across sessions and turns for model/prompt versions, retrieval, graph paths, tool calls, guardrail decisions, approvals, tokens, cost, errors, and component latency; provide redacted dashboards, SLO alerts, drift/anomaly monitoring, continuous evaluation, trace replay, audit lineage, and rollback signals.

REQUIRED SKILLS:

- Model Engineering : Python, PyTorch, Transformers, Hugging Face, TRL/PEFT, SFT, LoRA/QLoRA/DoRA, DPO/RLHF, BF16/FP16, DDP/FSDP/ DeepSpeed ZeRO, MLflow / W&B.;
- RAG & GraphRAG: Chunking, embeddings, hybrid search, reranking, Qdrant, Neo4j/AuraDB, Amazon Neptune, etc.; ontology, entity resolution, Cypher/Gremlin, vector-graph retrieval, grounding and citations.
- Inference & APIs: vLLM, TensorRT-LLM, SGLang, TGI/Triton, batching, KV/prefix cache, speculative decoding, quantization, FastAPI / OpenAI-compatible APIs, structured output and tool calling; intelligent model routing using rules, semantic/complexity classifiers, cascades, cost-quality-latency policies, FinOps budgets, metering, fallback, and routing observability.
- Platform & Resilience: Docker, Kubernetes, Helm, Terraform, CI/CD, Azure/AWS/GCP, multi-zone HA, cross-region DR, autoscaling, failover, backup/PITR, observability, RTO/RPO.
- Security & Quality: Layered guardrails, prompt-injection defense, PII/secrets, RBAC/ACL, HITL, auditability, OpenTelemetry tracing, evaluation, drift monitoring, latency/throughput/cost SLOs.

📌 Architect (Noida)
🏢 HCLTech
📍 Noida

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: architect (noida) / noida

Subscribe to this job alert:

Get the latest job offers by email for: architect (noida) / noida