We are looking for an Agentic AI – Infrastructure Intelligence Analyst to help transform traditional infrastructure operations from reactive support into predictive and proactive operations. The role will focus on analysing historical incident, request, problem, alert and infrastructure ticket data, identifying recurring patterns and operational trends, and working with infrastructure and AI teams to develop proactive remediation and Agentic AI use cases.
Key Responsibilities
Analyse historical ITSM ticket dumps across incidents, service requests, problems and changes.
Identify
Recurring incidents
High-volume ticket categories
Repetitive manual activities
Common root-cause patterns
Infrastructure failure trends
Seasonal and time-based patterns
Potential automation opportunities
Correlate ticket data with monitoring and infrastructure information.
Use Python, SQL and analytical techniques to identify operational trends and anomalies.
Create dashboards and reports highlighting:
Top incident drivers
Repeat issues
MTTR trends
Problem areas
Automation candidates
Proactive remediation opportunities
Work with infrastructure SMEs to convert identified patterns into actionable problem statements.
Support development of Agentic AI use cases for:
Incident triage
Root-cause analysis
Knowledge recommendation
Automated diagnostics
Predictive incident detection
Proactive remediation
Help build and maintain operational knowledge bases, SOPs, runbooks and remediation patterns for AI agents.
Validate AI recommendations against historical incidents and infrastructure standards.
Track outcomes and measure reduction in repeat incidents and manual effort.
Continuously identify new opportunities for AI-driven infrastructure operations.
Core Skills
Python
SQL
Data analysis
Machine Learning fundamentals
Statistical and trend analysis
Pattern recognition
Predictive analytics
Strong analytical and problem-solving skills
Infrastructure Knowledge – Positive to Have
Windows / Linux infrastructure
Cloud platforms such as Azure or AWS
Monitoring and observability concepts
ITSM / Incident / Problem / Change Management
Infrastructure automation
ServiceNow or equivalent ITSM platforms
Power BI or similar reporting tools
AI-based AI
Agentic AI concepts
Azure AI Foundry / Azure OpenAI / Amazon Bedrock
Python-based AI frameworks
Typical Use Cases
Ticket Intelligence
Analyse 6–12 months of incident data and identify the top recurring infrastructure problems.
Pattern & Trend Detection
Identify recurring CPU, memory, disk, network, backup or application-support incidents.
Problem Prediction
Detect infrastructure patterns that typically precede a major incident.
Proactive Operations
Recommend health checks or remediation before an incident is raised.
Automation Discovery
Identify repetitive tickets suitable for scripts, Ansible, workflows or autonomous agents.
Agentic Incident Management
Enable AI agents to analyse alerts, search historical tickets and knowledge, recommend root cause and propose remediation.
Success Measures
Reduction in repeat incidents
Reduction in manual ticket analysis
Increased automation opportunities identified
Improved problem-management effectiveness
Reduced MTTR
Increased percentage of incidents proactively prevented
- Increased reusable remediation patterns and knowledge articles