AI Agent Operations Engineer
Location: Pune(572)
Experience: 8–11 Years
Notice Period: Immediate–45 Days
Work Type: Client-facing production operations
Interview: 2 Technical Rounds + Client Round + HR
Preference: Local Pune candidates preferred
Role Overview
We are looking for an AI Agent Operations Engineer to lead production incident response, proactive monitoring, service reliability, and operational governance for AI-agent systems.
The role requires strong expertise in incident management, observability, SLOs, runbook operations, AI-agent failure analysis, and stakeholder communication.
Key Responsibilities
-
Act as Incident Commander during production incidents; lead bridges, escalations, communication, and resolution.
-
Conduct daily proactive trend reviews to identify degradation before alerts are triggered.
-
Coordinate rollback decisions for degraded releases and activate fallback processes when required.
-
Tune monitoring safeguards and alert thresholds to improve alert precision and reduce noise.
-
Lead post-incident reviews, track corrective actions, and maintain runbooks.
-
Analyze logs, traces, and monitoring data to identify production issues and trends.
-
Maintain shift-handover standards and mentor team members when required.
-
Communicate technical risks, incidents, and solutions clearly to non-technical stakeholders.
Must-Have Skills
-
Bachelor's/Master's degree in Computer Science or related field
-
relevant 6+ years in production operations
-
Experience in customer-facing production environments