21 Aug
|
Premium Aerotec
|
Bengaluru
21 Aug
Premium Aerotec
Bengaluru
Job Summary
Role: AIOps Leader
Date: August 2026
We are seeking a seasoned, hands-on AIOps Leader (7+ Years of Experience) to serve as the principal technical architect and lead builder for our centralized AI Product Support Operations function. Operating within a high-growth AI PSL (Product/Service Line), you will design, architect, and execute the end-to-end automation strategy that transforms raw operational chaos into scalable, self-healing, and data-driven infrastructure.
In this role, you will bridge software engineering, site reliability, and AI/ML architectures. You will lead the creation of intelligent diagnostic pipelines, custom RAG-driven knowledge tools, self-healing systems, and automated triage engines. You will work closely with cross-functional leadership, L1/L2 support teams, and platform engineering to systematically eliminate operational toil, optimize MTTR, and build proactive anomaly detection mechanisms across our AI ecosystem.
Number of positions: 1
Qualifications
- Education: Bachelor s or Master s degree in Computer Science, Software Engineering, Information Technology, or a related quantitative field.
- Overall Experience: 7+ years of hands-on experience across Software Engineering, Site Reliability Engineering (SRE), DevOps, or Systems Operations-with a focused concentration on cloud infrastructure automation and AI/ML operational tooling.
- Platform Specialization: At least 3+ years architecting and running operations directly within major cloud ecosystems (AWS or GCP), including native AI/ML compute platforms.
Responsibilities
- Architecture Advanced AI/ML Automation
- Self-Healing Infrastructure Workflows: Architect, build, and deploy auto-remediation routines, script-based diagnostic runners, and event-driven automation triggers that autonomously resolve platform issues.
- LLM RAG Systems Engineering: Design, implement,
and maintain advanced Retrieval-Augmented Generation (RAG) knowledge tools, vector databases, and LLM utilities that index telemetry, historic logs, and RCAs for instant incident context.
- Bot Agent Development: Lead the development and production rollout of conversational AI agents, custom webhooks, and self-service bots integrated into ticketing engines to automate Tier-1 and Tier-2 resolutions.
- Observability, Telemetry Predictive Analytics
- Observability Architecture: Build enterprise-grade telemetry ingestion workflows, automated log scraping, and context-enrichment pipelines that dynamically append system metrics directly to incident tickets upon creation.
- Predictive Anomaly Detection: Configure and tune real-time predictive alerting, log-pattern analysis, and AI-driven monitoring models across AWS, Azure, or GCP microservices.
- Operational BI Analytics: Architect and own centralized executive and operational dashboards (e.g., ServiceNow, Datadog) tracking MTTR velocity, system uptime, defect density, ticket deflection rates, and SLA/CSAT compliance.
- Operational Engineering L1/L2 Empowerment
- Toil Elimination: Continuously audit support operational bottlenecks across product teams, transforming high-volume manual intervention points into production-grade, single-click, or fully autonomous workflows.
- Log Metadata Standardization: Standardize system log outputs, stack-trace formatting,
and tagging taxonomy across all AI products to ensure platform telemetry remains machine-readable for AI engines.
- Technical Mentorship: Guide L1/L2 support engineers on best practices for automation, code-based triage, and log parsing.
Technical Essentials
- AWS GCP Native AI Platforms: Deep hands-on experience orchestrating production AI/ML workflows on AWS (Bedrock, SageMaker AI, OpenSearch, AWS Lambda) or GCP (Vertex AI, Vertex AI Agent Builder, Cloud Run, BigQuery).
- Cloud Infrastructure Infrastructure-as-Code (IaC): Advanced experience writing and managing cloud provisioning scripts using Terraform, AWS CloudFormation, or Google Cloud Deployment Manager to deploy auto-scaling, resilient operations environments.
- Containerization Orchestration: Production experience managing microservices via Kubernetes (EKS/GKE) and Docker to support agentic AI workers, vector indexing engines, and automated micro-tasks.
- Observability Cloud Telemetry: Proven capability to configure full-stack observability across cloud environments using AWS CloudWatch or GCP Cloud Logging/Monitoring to trigger automated alerts and log enrichment.
- Advanced Automation Scripting: Solid engineering capability in Python, Go, or Shell to build autonomous cloud functions (AWS Lambda/GCP Cloud Run), self-healing infrastructure scripts, and custom ITSM connectors (Jira/ServiceNow APIs).
- Enterprise Generative AI Stack: Production execution experience deploying RAG (Retrieval-Augmented Generation) architectures using cloud vector engines (Amazon Bedrock Knowledge Bases, GCP Vertex AI Search, Pinecone, or Qdrant) for automated incident context retrieval.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 AIOps Leader (Bengaluru)
🏢 Premium Aerotec
📍 Bengaluru