We are hiring for AI Model Architect.
Location: Noida/ Chennai/ Bangalore/ Hyderabad/ Pune
Experience: 10+ years
Key Responsibilities
AI Architecture & SLM Development
- Own enterprise GenAI architecture and build domain SLMs using Gemma, Llama, Qwen, Phi, Mistral, or equivalent models.
- Deliver multi-turn operational workflows for triage, diagnosis, approval-gated remediation, validation, ticket updates, escalation, durable state, resumability, and loop prevention.
Data Preparation & Ingestion
- Build event-driven and scheduled ingestion for PDF, DOCX, PPTX, XLSX, SharePoint, APIs, and approved email folders or attachments using change notifications, delta sync, checkpoints, and reconciliation.
- Implement idempotency, incremental reprocessing, OCR/parsing, deduplication, malware and PII/secret controls, metadata and ACL preservation, security-trimmed retrieval, durable queues, retries, dead-letter handling, lineage, freshness, and ingestion monitoring.
Fine-Tuning & Alignment
- Lead SFT/instruction tuning, continued pretraining, domain adaptation, LoRA, QLoRA, DoRA, DPO/RLHF, knowledge distillation, quantization-aware optimization, and model compression.
- Engineer training datasets with cleaning, labeling, synthetic data, JSONL/chat formats, versioning, quality gates, and train/validation/test splits.
RAG & Grounded AI
- Architect RAG with chunking, embeddings, hybrid dense/sparse search, metadata filtering, reranking, context assembly, citations, and grounded-ness evaluation using Qdrant or equivalent.
- Build GraphRAG with Neo4j/AuraDB, Amazon Neptune, etc.; apply ontologies, entity resolution, graph traversal, subgraph retrieval, Cypher/open Cypher or Gremlin, and vector-graph hybrid retrieval for multi-hop reasoning and provenance.
Inference Engineering & Model Serving
- Deliver low-latency serving with vLLM, TensorRT-LLM, SGLang, TGI, or Triton using energetic batching, paged attention, KV/prefix caching,
speculative decoding, streaming, routing, autoscaling, and INT4/INT8/FP8, AWQ, or GPTQ quantization.
- Develop intelligent model routing across SLMs, specialist models, and frontier LLMs using rules, semantic or complexity classification, model cascades, and learned routing; select the lowest-cost model that meets task-fit, quality, latency, context, privacy, safety, region, and availability requirements, with budgets, token metering, caching, circuit breakers, fallback, and continuous evaluation of routing accuracy, escalation rate, cost per successful task, and quality or latency regression.
- Declare and test TTFT, inter-token and p50/p95/p99 end-to-end latency, tokens/second, requests/second, concurrency, queue time, error rate, GPU/KV-cache utilization, context and token limits, cost, and performance under sustained, burst, failover, and degraded modes.
API, Platform & Production Operations
- Build OpenAI-compatible APIs with structured outputs, function/tool calling, orchestration, API gateway, authentication, RBAC/ACL, rate limits, quotas, secrets, and audit logging.
- Deploy through Docker, Kubernetes, Helm, Terraform, and CI/CD across Azure/AWS/GCP with multi-node/multi-zone HA, load balancing, autoscaling, automated failover, graceful degradation, blue-green/rolling release, rollback, cross-region DR, backups, point-in-time recovery, and tested RTO/RPO.
Evaluation, Security & Governance
- Define release gates for task success, Exact Match/F1, retrieval precision/recall, grounded-ness, citation coverage, hallucination/refusal, schema validity,
safety, latency, throughput, availability, RTO/RPO, and cost.
- Implement layered guardrails across input, retrieval, dialogue, routing, tool execution, and output: prompt-injection/jailbreak defense, PII/secret masking, grounding checks, tool allowlists, least privilege, HITL/dual approval, blast-radius and retry limits, kill switches, and deterministic fallback.
- Establish OpenTelemetry-compatible model/agent tracing across sessions and turns for model/prompt versions, retrieval, graph paths, tool calls, guardrail decisions, approvals, tokens, cost, errors, and component latency; provide redacted dashboards, SLO alerts, drift/anomaly monitoring, continuous evaluation, trace replay, audit lineage, and rollback signals.
REQUIRED SKILLS:
- Model Engineering: Python, PyTorch, Transformers, Hugging Face, TRL/PEFT, SFT, LoRA/QLoRA/DoRA, DPO/RLHF, BF16/FP16, DDP/FSDP/ DeepSpeed ZeRO, MLflow / W&B.;
- RAG & GraphRAG: Chunking, embeddings, hybrid search, reranking, Qdrant, Neo4j/AuraDB, Amazon Neptune, etc.; ontology, entity resolution, Cypher/Gremlin, vector-graph retrieval, grounding and citations.
- Inference & APIs: vLLM, TensorRT-LLM, SGLang, TGI/Triton, batching, KV/prefix cache, speculative decoding, quantization, FastAPI / OpenAI-compatible APIs, structured output and tool calling; intelligent model routing using rules, semantic/complexity classifiers, cascades, cost-quality-latency policies, FinOps budgets, metering, fallback, and routing observability.
- Platform & Resilience: Docker, Kubernetes, Helm, Terraform, CI/CD, Azure/AWS/GCP, multi-zone HA, cross-region DR, autoscaling, failover, backup/PITR, observability, RTO/RPO.
- Security & Quality: Layered guardrails, prompt-injection defense, PII/secrets, RBAC/ACL, HITL, auditability, OpenTelemetry tracing, evaluation, drift monitoring, latency/throughput/cost SLOs.
📌 Architect (Noida)
🏢 HCLTech
📍 Noida