05 Oct
|
Host360
|
Mumbai
Job Title: Pre-Sales - Generative & Agentic AI Solutions Architect
Company: Host360
Location: Onsite (Andheri East)
Employment Type: Full Time
Experience: 6–10 Years
Function: Solutions Architecture / Technical Pre-Sales
Focus: Generative AI, Agentic AI, RAG, LLM Inference & GPU-Accelerated Computing
CTC – As per Industry Standards
Role Overview
We are looking for a technically strong and customer-focused Generative & Agentic AI Solutions Architect – Pre-Sales to design, validate, and position enterprise AI solutions across Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), Agentic AI, and GPU-accelerated computing.
You will work closely with customers, Sales, Business Development, Engineering, and technology partners to understand business and technical requirements, develop end-to-end architectures, size AI infrastructure, build and validate PoCs, and guide customers from initial discovery through technical closure and production adoption.
The role requires a strong combination of AI technical depth, solution architecture, hands-on problem solving, and customer-facing pre-sales capability.
Key Responsibilities
- Engage directly with customers to understand their business challenges, AI use cases, data landscape, workloads, performance requirements, and existing infrastructure.
- Architect end-to-end Generative AI, LLM, RAG, and Agentic AI solutions aligned with customer requirements.
- Lead technical discovery sessions, architecture workshops, whiteboarding sessions, and solution-design discussions.
- Design RAG architectures covering data ingestion, chunking, embeddings, vector search, retrieval, reranking, context management, and enterprise data integration.
- Design Agentic AI architectures involving tool/function calling, orchestration, memory, planning, APIs, guardrails, human-in-the-loop, and enterprise system integration.
- Design and size scalable LLM inference platforms and GPU-accelerated infrastructure based on model, workload, concurrency, context length, latency, throughput, and availability requirements.
- Analyze end-to-end AI workloads and recommend approaches to improve latency, throughput, scalability, GPU utilization, reliability, and cost efficiency.
- Evaluate and recommend appropriate models, inference engines, AI frameworks, vector databases, and deployment architectures.
- Build or lead PoCs, technical demonstrations, benchmarks, and reference solutions to validate proposed architectures and customer use cases.
- Support Sales and Business Development teams throughout the technical sales cycle, including customer presentations, solution positioning, demonstrations, technical qualification, and objection handling.
- Translate customer requirements into solution architectures, infrastructure sizing, BOMs, technical proposals, RFP/RFI responses, and implementation scope/SOWs.
- Identify technical risks, dependencies, assumptions, security requirements, and integration challenges before solution commitment.
- Ensure proposed solutions are technically feasible, scalable, secure, supportable, and commercially practical.
- Collaborate with Engineering and Delivery teams to ensure effective transition from pre-sales to implementation and production deployment.
- Collaborate with NVIDIA and other technology partners for architecture validation, GPU sizing, performance optimization, product feedback, and complex customer requirements.
- Provide technical leadership and guidance on best practices for production-grade Generative AI and Agentic AI deployments.
- Work with Enterprise, Government, Research, and Public Sector customers on strategic AI initiatives.
- Stay current with developments in LLMs, Agentic AI, RAG, inference engines, GPU technologies, and distributed AI systems.
Must-Have Skills & Experience
- 6–10 years of overall experience across Solution Architecture, Technical Pre-Sales, AI/ML, Cloud, Data Platforms, Infrastructure, or related technologies.
- Strong customer-facing experience in a Solution Architect, Pre-Sales Architect, AI Architect, or Technical Consultant role.
- Strong understanding of Generative AI, LLMs, RAG, and Agentic AI architectures.
- Hands-on experience designing or delivering LLM, RAG, or Agentic AI solutions in production or production-like environments.
- Ability to translate business requirements into end-to-end technical architectures and solution designs.
- Practical understanding of RAG components including embeddings, vector databases, semantic/hybrid search, retrieval, reranking, context management, and evaluation.
- Understanding of Agentic AI concepts including tool/function calling, orchestration, memory/state, planning, multi-step workflows, guardrails, and human-in-the-loop patterns.
- Strong understanding of LLM inference concepts including:
1. Latency and throughput
2. Concurrency and batching
3. GPU memory and utilization
4. KV cache
5. Quantization
6. Multi-GPU deployment
7. Scalability and cost optimization
- Experience with Python, Linux, Docker, and Kubernetes.
- Experience with AI/ML ecosystems such as PyTorch and Hugging Face.
- Experience or exposure to LLM serving technologies such as NVIDIA NIM, TensorRT-LLM, Triton Inference Server, vLLM, or equivalent.
- Working understanding of GPU infrastructure, model sizing, networking, storage, security, and observability requirements for AI workloads.
- Experience conducting customer discovery, architecture workshops, technical presentations, demonstrations, and PoCs.
- Experience supporting technical pre-sales activities including solution sizing, RFP/RFI responses, technical proposals, and technical closure.
- Strong written and verbal communication skills with the ability to engage developers, architects, infrastructure teams, business stakeholders, and senior customer leadership.
Preferred / Positive-to-Have Skills
- Experience with NVIDIA AI Enterprise, NIM, NeMo, TensorRT-LLM, Triton Inference Server, Dynamo, or other NVIDIA AI technologies.
- Experience designing or operating GPU clusters and distributed AI infrastructure.
- Understanding of advanced inference architectures including distributed/disaggregated serving, intelligent routing and scheduling, KV-cache optimization, heterogeneous inference, and model parallelism.
- Experience with LangGraph, LlamaIndex, LangChain, Semantic Kernel, or similar Agentic AI frameworks.
- Experience with model customization, LoRA/PEFT, fine-tuning, quantization, and model evaluation.
- Experience with Kubernetes GPU scheduling and production AI platform architecture.
- Experience designing AI solutions across on-premises, private cloud, and public cloud environments.
- Understanding of AI security, governance, observability, responsible AI, and guardrail frameworks.
- Experience with GPU/infrastructure sizing, BOM preparation, capacity planning, and performance benchmarking.
- Experience with high-performance networking and storage for AI workloads.
- Experience working with Government, Research, Public Sector, CSP, or large enterprise customers.
The ideal candidate combines solid customer engagement and pre-sales capability with sufficient hands-on technical depth to design, validate, defend, and optimize enterprise AI solutions.
📌 Pre-Sales - Generative & Agentic AI Solutions Architect (Mumbai)
🏢 Host360
📍 Mumbai