14 Sep
|
Promaynov Advisory Services
|
Bengaluru
14 Sep
Promaynov Advisory Services
Bengaluru
Please find the below details :
Role
Consultant – AI Application Support (Level 2) Experience :
4-6 years Location :
Bangalore & Gurgaon
Role purpose
Provide advanced
Level 2 (L2)
support for AI-based applications by performing deep triage, troubleshooting, and resolution of incidents and service requests escalated from
L1 Service Desk . Distinguish functional queries from technical incidents , restore service quickly, and escalate to
L3 Product Support (Advisory UK service lines)
with complete diagnostic context. Strengthen operational readiness for rapid scale (20 apps → 100+ apps) through knowledge management ,
documentation , and observability/reliability improvements.
Key responsibilities (what you will do) 1) L2 Ticket Handling, Triage & Ownership (ServiceNow) Pick up escalations from
L1 Service Desk and take ownership through resolution or appropriate escalation. Validate and confirm impact, urgency, and priority , ensuring correct categorization and assignment . Perform structured triage to classify issues as: Functional queries
(usage/how-to, expected behaviour, configuration guidance), or Technical incidents
(errors, outages, latency, integration failures, platform degradation). Apply runbooks/diagnostic playbooks to resolve issues within L2 scope and document all actions in ServiceNow. 2) Technical Troubleshooting Across AI + Cloud Stack Troubleshoot failures spanning: AI platforms : OpenAI / Azure AI patterns such as authentication errors, throttling (rate limits), token/context issues, deployment/model availability, safety/content filtering impacts, latency/timeouts. Cloud services
(Azure preferred; AWS/GCP familiarity): identity/IAM, networking, certificates, secrets/key vaults, API gateways, resource limits, configuration drift.
Application dependencies : integrations, data sources, retrieval pipelines (if applicable), middleware, and downstream service availability.
3) High-Quality Escalation to L3 (Advisory UK Product Support) When L2 resolution is not possible, escalate to L3 with a complete evidence package to minimize back-and-forth:
reproduction steps and expected vs actual behaviour timestamps, request IDs/correlation IDs logs/metrics/traces (e.g., App Insights/Azure Monitor where applicable) environment details (prod/non-prod, region, model/deployment name, endpoint) recent changes/releases, interim mitigations applied, and L2 hypothesis sanitized payload samples (compliant with data handling rules) 4) Knowledge Management & Documentation (Deflection + Speed-to-Resolve) Create and maintain
Support Manuals ,
Runbooks ,
SOPs , and
Troubleshooting Guides for each supported AI application. Write and curate
Knowledge Articles in ServiceNow to enable L1 deflection and consistent resolution. Maintain "Known Errors / Known Issues" records and update them based on recurring patterns and fixes.
5) Observability, Reliability & Operational Readiness Identify monitoring gaps and recommend improvements (alerts/dashboards/logging standards) to reduce MTTR. Support go-live readiness for new AI apps by ensuring: support model and escalation paths are defined runbooks and known dependencies are documented logging/telemetry is sufficient for L2 diagnosis common failure modes and mitigations are captured 6) Problem Management & Continuous Improvement Detect trends from incident patterns (e.g., repeated throttling, recurring auth failures, dependency outages). Raise and contribute to problem records with evidence and recommended preventive actions. Participate in post-incident reviews and implement corrective/preventive tasks within the support scope.
Skills required (must-have) Working knowledge of
AI platforms : OpenAI / Azure AI, common operational failure patterns (auth, throttling, safety filters,
latency, deployment errors). Robust cloud fundamentals :
Azure
(preferred) plus familiarity with
AWS and GCP
(networking basics, IAM, logging/monitoring, secrets management, API services). Strong problem-solving and triage skills; ability to isolate whether issue is app defect, platform issue, integration issue, or user/process issue. ServiceNow
(or equivalent ITSM) proficiency: incident/request/problem workflows, categorization, SLAs, knowledge base. Excellent documentation skills: clear, structured runbooks and KB articles. Strong written and verbal communication; comfortable collaborating with L1, app owners, platform teams, and L3 engineering/product support.
Experience (typical) 4+ years in production support / operations (L2/L3-facing), ideally in cloud-native environments. Exposure to AI/ML or AI-enabled applications in production is strongly preferred.
Desirable / nice-to-have Familiarity with observability tooling and queries (e.g., KQL/App Insights, Azure Monitor; equivalents in AWS/GCP). Basic scripting/automation (PowerShell/Python) to accelerate diagnostics and reduce repeat toil. ITIL Foundation or equivalent service management training.
Key measures of success (KPIs) Time-to-triage from L1 escalation MTTR for L2-resolvable incidents SLA compliance
(response/resolution) for assigned queue/categories First-time-right escalation quality
(L3 acceptance rate; reduced ping-pong) Knowledge contribution rate
(KAs/runbooks created/updated; KB reuse/deflection) Recurring incident reduction through problem management inputs
Interfaces & working relationships L1 Service Desk:
coaching via KB/runbooks; quality feedback on triage and ticket capture L3 Product Support (Advisory UK):
escalation partner for deep product defects and complex technical issues Cloud/Platform teams:
support for platform incidents, capacity/limits, identity/networking issues App Owners / SMEs:
functional clarifications, change/release coordination, operational readiness
📌 AI Application Support (Bengaluru)
🏢 Promaynov Advisory Services
📍 Bengaluru