Our client's monitoring tools turn a single failure into dozens of separate tickets. We're building an AI incident-intelligence layer over their Zabbix and ServiceNow stack that:
• Correlates alerts into one parent incident
• Drafts root-cause analysis automatically
• Automates access revocation for leavers
All of this is built without replacing any tool they already own.
What You'll Do
• Integrate Zabbix and other monitoring feeds into ServiceNow using webhooks, REST and syslog adapters
• Configure ServiceNow Event Management correlation rules and maintain CMDB service relationships (cmdb_rel_ci)
• Build and run n8n workflows for deduplication, alert suppression, ticket creation and self-healing playbooks
• Build a fail-safe path so alerts reach ServiceNow directly if the correlation layer goes down
• Deploy the RAG pipeline that drafts root-cause analysis from historical incident notes
• Automate Joiner-Mover-Leaver (JML) access revocation across Okta / Entra ID and downstream apps
• Implement approval gates in Teams or Slack for any destructive automated action
• Measure ticket-noise reduction, MTTI and false-positive rates using real client data
• Serve as the daily technical point of contact for the client's SRE, ServiceNow and IAM teams, and share a daily progress update
Must-Have Skills
ServiceNow depth and integration experience are non-negotiable. The AI layer can be learned on the job if the rest is solid.
• ServiceNow: 2+ years hands-on with ITOM or Event Management, the CMDB data model (cmdb_ci, cmdb_rel_ci), Table API and Scripted REST APIs. You have built alert correlation or event rules, not just used them.
• Monitoring: Production experience with Zabbix, plus at least one of Datadog, Dynatrace, Splunk or ELK (triggers, webhooks, alert payloads)
• Integration: REST, webhooks, OAuth2 and JSON schema mapping. Comfortable reading a vendor API doc and shipping an adapter in a day.
• Python: Production scripting and services (Node.js is a plus)
• Workfl