Pre-Sales - Generative & Agentic AI Solutions Architect (Mumbai)

Pre-Sales - Generative & Agentic AI Solutions Architect (Mumbai)

05 Oct
|
Host360
|
Mumbai

05 Oct

Host360

Mumbai

Job Title: Pre-Sales - Generative & Agentic AI Solutions Architect

Company: Host360

Location: Onsite (Andheri East)

Employment Type: Full Time

Experience: 6–10 Years

Function: Solutions Architecture / Technical Pre-Sales

Focus: Generative AI, Agentic AI, RAG, LLM Inference & GPU-Accelerated Computing

CTC – As per Industry Standards

Role Overview

We are looking for a technically strong and customer-focused Generative & Agentic AI Solutions Architect – Pre-Sales to design, validate, and position enterprise AI solutions across Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), Agentic AI, and GPU-accelerated computing.

You will work closely with customers, Sales, Business Development, Engineering, and technology partners to understand business and technical requirements, develop end-to-end architectures, size AI infrastructure, build and validate PoCs, and guide customers from initial discovery through technical closure and production adoption.

The role requires a strong combination of AI technical depth, solution architecture, hands-on problem solving, and customer-facing pre-sales capability.

Key Responsibilities

- Engage directly with customers to understand their business challenges, AI use cases, data landscape, workloads, performance requirements, and existing infrastructure.
- Architect end-to-end Generative AI, LLM, RAG, and Agentic AI solutions aligned with customer requirements.
- Lead technical discovery sessions, architecture workshops, whiteboarding sessions, and solution-design discussions.
- Design RAG architectures covering data ingestion, chunking, embeddings, vector search, retrieval, reranking, context management, and enterprise data integration.
- Design Agentic AI architectures involving tool/function calling, orchestration, memory, planning, APIs, guardrails, human-in-the-loop, and enterprise system integration.
- Design and size scalable LLM inference platforms and GPU-accelerated infrastructure based on model, workload, concurrency, context length, latency, throughput, and availability requirements.
- Analyze end-to-end AI workloads and recommend approaches to improve latency, throughput, scalability, GPU utilization, reliability, and cost efficiency.
- Evaluate and recommend appropriate models, inference engines, AI frameworks, vector databases, and deployment architectures.
- Build or lead PoCs, technical demonstrations, benchmarks, and reference solutions to validate proposed architectures and customer use cases.




- Support Sales and Business Development teams throughout the technical sales cycle, including customer presentations, solution positioning, demonstrations, technical qualification, and objection handling.
- Translate customer requirements into solution architectures, infrastructure sizing, BOMs, technical proposals, RFP/RFI responses, and implementation scope/SOWs.
- Identify technical risks, dependencies, assumptions, security requirements, and integration challenges before solution commitment.
- Ensure proposed solutions are technically feasible, scalable, secure, supportable, and commercially practical.
- Collaborate with Engineering and Delivery teams to ensure effective transition from pre-sales to implementation and production deployment.
- Collaborate with NVIDIA and other technology partners for architecture validation, GPU sizing, performance optimization, product feedback, and complex customer requirements.
- Provide technical leadership and guidance on best practices for production-grade Generative AI and Agentic AI deployments.
- Work with Enterprise, Government, Research, and Public Sector customers on strategic AI initiatives.
- Stay current with developments in LLMs, Agentic AI, RAG, inference engines, GPU technologies, and distributed AI systems.

Must-Have Skills & Experience

- 6–10 years of overall experience across Solution Architecture, Technical Pre-Sales, AI/ML, Cloud, Data Platforms, Infrastructure, or related technologies.
- Strong customer-facing experience in a Solution Architect, Pre-Sales Architect, AI Architect, or Technical Consultant role.
- Strong understanding of Generative AI, LLMs, RAG, and Agentic AI architectures.
- Hands-on experience designing or delivering LLM, RAG, or Agentic AI solutions in production or production-like environments.
- Ability to translate business requirements into end-to-end technical architectures and solution designs.
- Practical understanding of RAG components including embeddings, vector databases, semantic/hybrid search, retrieval, reranking, context management, and evaluation.
- Understanding of Agentic AI concepts including tool/function calling, orchestration, memory/state, planning, multi-step workflows, guardrails, and human-in-the-loop patterns.




- Strong understanding of LLM inference concepts including:

1. Latency and throughput
2. Concurrency and batching
3. GPU memory and utilization
4. KV cache
5. Quantization
6. Multi-GPU deployment
7. Scalability and cost optimization

- Experience with Python, Linux, Docker, and Kubernetes.
- Experience with AI/ML ecosystems such as PyTorch and Hugging Face.
- Experience or exposure to LLM serving technologies such as NVIDIA NIM, TensorRT-LLM, Triton Inference Server, vLLM, or equivalent.
- Working understanding of GPU infrastructure, model sizing, networking, storage, security, and observability requirements for AI workloads.
- Experience conducting customer discovery, architecture workshops, technical presentations, demonstrations, and PoCs.
- Experience supporting technical pre-sales activities including solution sizing, RFP/RFI responses, technical proposals, and technical closure.
- Strong written and verbal communication skills with the ability to engage developers, architects, infrastructure teams, business stakeholders, and senior customer leadership.

Preferred / Positive-to-Have Skills

- Experience with NVIDIA AI Enterprise, NIM, NeMo, TensorRT-LLM, Triton Inference Server, Dynamo, or other NVIDIA AI technologies.
- Experience designing or operating GPU clusters and distributed AI infrastructure.
- Understanding of advanced inference architectures including distributed/disaggregated serving, intelligent routing and scheduling, KV-cache optimization, heterogeneous inference, and model parallelism.
- Experience with LangGraph, LlamaIndex, LangChain, Semantic Kernel, or similar Agentic AI frameworks.
- Experience with model customization, LoRA/PEFT, fine-tuning, quantization, and model evaluation.
- Experience with Kubernetes GPU scheduling and production AI platform architecture.
- Experience designing AI solutions across on-premises, private cloud, and public cloud environments.
- Understanding of AI security, governance, observability, responsible AI, and guardrail frameworks.
- Experience with GPU/infrastructure sizing, BOM preparation, capacity planning, and performance benchmarking.
- Experience with high-performance networking and storage for AI workloads.
- Experience working with Government, Research, Public Sector, CSP, or large enterprise customers.

The ideal candidate combines solid customer engagement and pre-sales capability with sufficient hands-on technical depth to design, validate, defend, and optimize enterprise AI solutions.

📌 Pre-Sales - Generative & Agentic AI Solutions Architect (Mumbai)
🏢 Host360
📍 Mumbai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: pre-sales - generative & agentic ai solutions architect (mumbai) / mumbai

Subscribe to this job alert:

Get the latest job offers by email for: pre-sales - generative & agentic ai solutions architect (mumbai) / mumbai