07 Aug
|
GLIDER.ai
|
Bengaluru
07 Aug
GLIDER.ai
Bengaluru
Objective
Work with the AI architecture and engineering teams to build and operationalize the cloud foundation required for agentic AI applications — enabling secure, repeatable, and automated deployment across the Marketing AI stack.
Key Responsibilities
- Implement GitOps pipelines for AI-agent deployment, working with the Data Office and engineering teams.
- Review and configure cloud (Azure) subscriptions and landing zones — networking, identity, secrets management, monitoring, and deployment automation — in line with HPE cybersecurity requirements, and support Architecture Review Board (ARB) requests.
- Create reusable infrastructure templates, deployment pipelines, and operational runbooks for all AI applications.
- Support onboarding of the agent platform, Databricks, and inter-agent infrastructure components (MCP, agent-to-agent services) into production environments.
- Execute and document deployment, security, performance, and operational-readiness testing.
AgentOps & Production Operations
- Stand up monitoring, observability, and centralized logging for AI agents and services, with actionable alerting on health, performance, and cost.
- Own agent lifecycle management in production — automated deployment, versioning, controlled rollout, rollback, and upgrades.
- Implement runtime governance and guardrails (usage controls, safety and policy checks, human-in-the-loop hooks) and ensure agent actions are auditable.
- Define operational support processes and runbooks so AI systems run reliably at enterprise scale — not just get built.
Expected Deliverables
- A working GitOps deployment process for AI agents.
- A secure cloud landing zone and shared infrastructure services.
- Standardized deployment pipelines for AI applications.
- Operational documentation and automated deployment workflows.
- An operational AgentOps setup — monitoring, logging, and alerting across the AI applications.
- Automated agent deployment, versioning, and rollback processes, with runtime guardrails, audit logging, and production runbooks.
Required Skills & Experience
- Proven cloud infrastructure / DevOps engineering experience (typically 10+ years), ideally on Microsoft Azure — subscriptions, landing zones, networking, identity (Entra ID), Key Vault / secrets management, and monitoring/observability.
- Infrastructure-as-Code (e.g., Terraform, Bicep, or ARM) and GitOps / CI-CD tooling for automated, repeatable deployments.
- Containerization and orchestration (Docker, Kubernetes) applied to application and agent deployment.
- Working knowledge of Databricks and AI/agent infrastructure concepts.
- Robust security, compliance, and enterprise-governance mindset; comfortable working within formal architecture-review processes.
- Bachelor's or master's degree in computer science, Software Engineering, or a related field (or equivalent experience).
- Hands-on AgentOps / MLOps / LLMOps — production monitoring and observability (metrics, tracing, logging), agent and model lifecycle management, and CI/CD for AI workloads.
- Proven experience operating AI/LLM systems reliably in production: guardrails, evaluation, incident response, and cost/performance optimization.
Nice to Have
- Experience deploying LLM / generative-AI and multi-agent workloads to production.
- Cloud certifications (Azure preferred; AWS/GCP a plus).
📌 AI Cloud Architect (Bengaluru)
🏢 GLIDER.ai
📍 Bengaluru