22 Sep
|
Tata Consultancy Services
|
India
22 Sep
Tata Consultancy Services
India
Experience:
10+ years of experience in software/AI engineering, with strong hands-on experience designing and implementing Generative AI, Agentic AI, and Multi-Agent Systems on Google Cloud Platform.
Agentic AI & Multi-Agent Systems:
Strong experience with agent orchestration, planning, reasoning, tool/function calling, memory, workflow management, Human-in-the-Loop, and agent-to-agent communication. Hands-on exposure to Vertex AI Agent Builder, ADK, Agent Engine, MCP, A2A, or similar frameworks.
Generative AI & RAG:
Robust experience with LLMs, Gemini, prompt engineering, RAG, embeddings, vector search/vector databases, grounding, reranking, and context management.
Python & AI Application Development:
Advanced Python programming with experience in FastAPI, REST APIs, SDKs, asynchronous processing, and development of production-grade AI applications and services.
Google Cloud AI Platform:
Strong hands-on experience with Vertex AI, Gemini APIs, BigQuery, Cloud Storage, IAM, Secret Manager, Cloud Run, and/or GKE for building scalable AI solutions.
AI Security & Governance:
Experience with IAM, secure authentication/authorization, data protection, prompt-injection controls, guardrails, API security, and Responsible AI practices.
Agent Evaluation & Observability:
Experience evaluating response quality, groundedness, hallucination, retrieval accuracy, tool-call success, task completion, latency, and cost, along with tracing, logging, monitoring, and production troubleshooting.
AI Deployment & LLMOps:
Hands-on experience with Docker, Kubernetes, CI/CD, Terraform, automated deployments, monitoring, scalability, and production support for AI workloads.
Enterprise Integration:
Ability to integrate AI agents with enterprise applications, databases, APIs, SaaS platforms, and external tools using secure integration patterns.
Collaboration & Solution Design:
Strong problem-solving and communication skills with the ability to translate business requirements into scalable, secure, and production-ready AI solutions.
Architecture & Ownership:
Ability to independently design scalable, reliable and secure multi-agent solutions, translate business requirements into technical architecture, manage ambiguity and collaborate with architecture, security, DevOps and application teams.
Key Responsibilities
- Understand business problems, stakeholder expectations, functional requirements, non-functional requirements, security constraints, and success criteria, and translate them into scalable Generative AI and Agentic AI solution designs.
- Define the end-to-end AI solution architecture, including LLM selection, agent architecture, RAG design, enterprise integrations, APIs, data flow, security, deployment model, observability, and infrastructure components.
- Evaluate technical feasibility and recommend appropriate build-versus-buy, model, framework, architecture, and cloud-service choices based on business value, scalability, security, performance, and cost.
- Design and develop production-grade Generative AI, Agentic AI, Multi-Agent, and RAG solutions using appropriate frameworks, models, tools, and Google Cloud services.
- Lead integration of AI solutions with enterprise applications, databases, APIs, SaaS platforms, business workflows, and external tools,
ensuring secure and reliable communication.
- Define and implement AI evaluation and testing strategies, covering functional testing, agent behavior, retrieval quality, groundedness, hallucination, tool-call accuracy, task completion, regression testing, performance, and security testing.
- Establish appropriate Human-in-the-Loop, guardrails, approval workflows, fallback mechanisms, and error-handling patterns for critical or sensitive AI use cases.
- Ensure AI solutions follow enterprise security, privacy, Responsible AI, IAM, data-protection, and governance standards, including protection against prompt injection, unauthorized data access, and unsafe tool execution.
- Design and implement CI/CD and LLMOps practices for automated testing, versioning, deployment, configuration management, rollback, and release of prompts, agents, APIs, and AI applications.
- Deploy and operate AI workloads on Vertex AI, Cloud Run, GKE, or other approved GCP services, ensuring scalability, high availability, resilience, and production readiness.
- Continuously optimize LLM usage, token consumption, infrastructure utilization, vector retrieval, caching, model selection, concurrency, and API calls to improve solution performance and reduce operational cost.
- Implement comprehensive logging, tracing, monitoring, alerting, and observability to support troubleshooting complex production issues across agents, LLMs, RAG pipelines, APIs, cloud infrastructure, integrations, security, and performance, and drive root-cause analysis and corrective actions.
- Define and promote AI engineering best practices, reusable architecture patterns, coding standards, evaluation standards, documentation standards, and governance controls across projects.
📌 Multi-Agent Systems (MAS) using frameworks like the Agent Development (India)
🏢 Tata Consultancy Services
📍 India