19 Sep
|
Straive
|
Bengaluru
Principal AI Architect — Generative AI & Enterprise Delivery
India (Hybrid) | Data, Analytics & AI | Full-time | Hands-on Architecture & Senior Engineering
About Us
Straive is a global leader in data analytics and AI operationalization, helping enterprises embed advanced AI and data capabilities into core business workflows to deliver measurable business outcomes and ROI. With a workforce of ~20,000 professionals serving 350+ clients across 30+ markets, Straive combines technical scale with deep domain expertise. A key differentiator is its network of over 6,000 subject matter experts who specialize in managing and enriching complex, unstructured data enabling organizations to build AI systems grounded in accuracy, context, and business relevance.
The company continues to earn recognition from leading industry analysts and was recently named a Leader in AIM's 2026 Generative AI and Data Engineering PeMa Quadrants. Backed by EQT—a purpose-driven global investment organization which was ranked among the world's leading private equity firms by PEI in 2025—Straive is positioned as a high-value alternative to traditional IT services providers, combining domain-led intelligence with AI execution at scale.
About the Role
Straive is seeking a hands-on Principal AI Architect & Senior Developer to serve as the technical authority for our enterprise Generative AI delivery practice. In this role, you will lead the end-to-end architectural design, backend engineering, and production deployment of enterprise Generative AI platforms, multi-agent systems, document intelligence pipelines, and real-time streaming LLM applications.
This is an intensely hands-on role requiring expert backend proficiency in Python and SQL alongside deep expertise in Generative AI architectures. You will move beyond simple API orchestration to fine-tune open-weight models, build custom agentic workflows, engineer advanced multimodal RAG pipelines, write high-concurrency streaming backend microservices (FastAPI/AsyncIO), enforce AI security/PII guardrails, and optimize LLM unit economics (inference cost, latency, accuracy) for production deployments.
Key Responsibilities
- Hands-on Gen AI Architecture & System Design: Architect scalable enterprise Gen AI solutions—including multi-agent orchestration (LangGraph or AutoGen or CrewAI), advanced RAG retrieval pipelines, real-time streaming interfaces (SSE, WebSockets), prompt security, guardrails, and automated evaluation suites.
- High-Concurrency Backend Engineering & SQL: Write clean, async,
high-performance Python backend code (FastAPI or AsyncIO or PyDantic) and write complex, optimized SQL queries to build robust microservices, caching layers (Redis), queueing systems (RabbitMQ/Celery), and vector search integrations.
- Open-Weight Fine-Tuning & AI Economics: Own open-weight model fine-tuning (Llama or Mistral or Qwen or Phi) using LoRA, QLoRA, and full tuning. Drive inference optimization via quantization (AWQ, GPTQ), model distillation, prompt caching, and high-throughput GPU serving (vLLM/Triton).
- Agentic Workflows & Tool Integration (MCP): Architect autonomous multi-agent workflows, planner-worker execution patterns, and custom Model Context Protocol (MCP) servers connecting LLMs securely to enterprise databases, APIs, and legacy systems of record.
- Multimodal RAG & Document Intelligence: Build complex document intelligence and multimodal ingestion pipelines to parse, OCR, chunk, and embed structured/unstructured documents (PDFs, tables, images) using Textract, Unstructured, and Vision-Language Models (VLMs).
- LLMOps, Telemetry, Guardrails & Governance: Establish end-to-end LLMOps infrastructure—automated CI/CD eval gates (Ragas, TruLens), hallucination detection, telemetry/tracing (LangSmith, OpenTelemetry, Arize), and AI safety controls (NeMo Guardrails, Llama Guard, PII redaction).
- Technical Leadership, RFPs & Client Advisory: Lead and mentor 30+ AI engineers and developers through architecture reviews and hands-on code reviews. Partner with client CTOs, CDOs, and commercial teams on RFPs, technical proposals, discovery workshops, and effort estimations.
Required Experience and Qualifications
- 12+ years of skilled software engineering experience across system design, distributed backend development, and AI/ML, with a minimum of 4+ years dedicated to hands-on Generative AI and LLM application architecture.
- Expert-level hands-on proficiency in Python (FastAPI or AsyncIO or PyDantic or Celery) and SQL (PostgreSQL/pgvector, MySQL, complex joins, CTEs, query optimization) for building production microservices.
- Proven track record architecting and shipping production Gen AI systems using multi-agent frameworks (LangGraph or AutoGen or CrewAI or LangChain)
and advanced RAG architectures (hybrid semantic + BM25 retrieval/re-ranking models).
- Hands-on experience fine-tuning open-weight LLMs/SLMs (Llama/Mistral/Qwen/Phi), dataset curation, LoRA/QLoRA execution, and serving models under production load using vLLM, Triton, or Ollama.
- Strong practical depth with Vector Databases (OpenSearch Vector Search/ Qdrant/ Milvus/ FAISS/ Pinecone) and document parsing/multimodal pipelines (AWS Textract, OCR, Vision-Language Models).
- Demonstrated ability to optimize LLM inference cost and latency via quantization (AWQ/GPTQ, GGUF), prompt engineering/caching, model routing, and batching.
- Hands-on expertise in LLMOps telemetry and AI security: evaluation frameworks (Ragas/TruLens), tracing (LangSmith/OpenTelemetry), prompt injection protection, PII masking, and guardrails (NeMo Guardrails/Llama Guard).
- Solid command of containerization (Docker, Kubernetes), streaming APIs (SSE, WebSockets), GPU serving, and cloud AI platforms across either AWS (Bedrock/SageMaker), Azure (Azure OpenAI), or GCP (Vertex AI).
Preferred Experience & Education
- Hands-on exposure to Model Context Protocol (MCP) servers, Claude Skills, and agentic tool-use protocols.
- Working command of enterprise AI risk and compliance frameworks.
- Familiarity with enterprise data platforms (Databricks MLflow/Snowflake) or classical ML (XGBoost, scikit-learn).
- Bachelor's or Master's degree in Computer Science, Software Engineering, Artificial Intelligence, or a related quantitative field from Tier 1/Tier 2 colleges)
Recognition & Achievements
- We have been recognized as a Star Performer in Data & AI Services Specialists - Everest's North America PEAK Matrix 2025, and as a Leader in AIM's PeMa Quadrant of Agentic AI Service Providers - 2025.
- In Nov 2023, Straive acquired Gramener, an award-winning, design-led data science company, enhancing our data, analytics, and AI capabilities. In June 2025, we acquired SG Analytics, a leading provider of AI-powered insights and contextual analytics services.
For more information, please visit our Website: https://www.straive.com/
Equal Opportunity Employer
Straive is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All employment decisions are based on business needs, job requirements, and individual qualifications, without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran, or disability status.
📌 Principal Architect (Bengaluru)
🏢 Straive
📍 Bengaluru