18 Aug
|
Accenture
|
India
Project Role : Operations Engineer
Project Role Description : Support the operations and/or manage delivery for production systems and services based on operational requirements and service agreement.
Must have skills : Agentic Orchestration, Agentic AI, AI Orchestration &LLM;
Positive to have skills : NA
Minimum 3 year(s) of experience is required
Educational Qualification : 15 years full time education
Summary:
A Tools & Platforms Service Engineer is responsible for the day-to-day operation, monitoring, support, and maintenance of the infrastructure engineering tooling estate — spanning observability platforms, infrastructure-as-code tooling, CI/CD pipelines, ITSM platforms, internal developer portals, secret management, and AI-augmented operations tooling. The role ensures that the platforms enabling Cloud, Network, Security, Database, and Voice teams to operate at scale remain available, performant, and fit for purpose.
At Level 9 / 10, this individual forms the primary operational execution layer within the tools and platforms tower — a proficient hands-on engineer who keeps platform tooling running reliably, resolves incidents and service requests efficiently, executes changes cleanly, and supports the adoption of AI-augmented operational capabilities across the infrastructure engineering function.
PLATFORMS IN SCOPE
Observability:
– ELK/Splunk
– OpenTelemetry
– Alertmanager
Automation & IaC:
– Terraform (consumer)
– Ansible playbooks
– GitHub Actions
– HashiCorp Vault
ITSM & DevOps:
– ServiceNow
– Jira/Confluence
– Xmatters
– CMDB tooling
AI-Augmented Ops:
– LLM runbook support
– AI alert triage
– Chatbot ops interfaces
– Prompt execution
– LLMOps monitoring
Roles & Responsibilities:
Observability & Monitoring Platform Operations
– Monitor and maintain observability platforms — Splunk/ELK — ensuring dashboards, alert rules, and data pipelines remain operational and accurate across all infrastructure tiers
– Triage and resolve observability platform incidents — broken scrapers, missing metrics, failed log ingestion, dashboard errors — within defined SLA commitments
– Manage alert configuration — creating, updating, and tuning alert rules in Alertmanager, or equivalent to reduce noise and maintain actionable alerting across infrastructure teams
– Operate ELK stack or Splunk log pipelines — managing index lifecycle, ingest pipeline health, and log source connectivity for infrastructure platforms
– Execute observability platform changes — dashboard deployments, alert rule updates,
retention policy changes — following approved change management procedures
– Maintain runbooks and operational documentation for all observability platforms — keeping procedures current, accurate, and audit-ready
ITSM, DevOps Tooling & AI-Augmented Operations
– Manage Xmatters — maintaining escalation policies, on-call schedules, alert routing rules, and integration health with upstream observability platforms
– Support Jira and Confluence platform operations — managing project configurations, workflow rules, automation triggers, and user access across engineering teams
– Monitor and support AI-augmented operations tooling — LLM runbook automation pipelines, agentic ITSM workflow health, and AI alert correlation services — ensuring AI tools are operational and escalating failures to Principal Engineers
– Support engineering teams using AI-augmented operations capabilities — triaging issues with LLM-powered runbook execution, RAG knowledge base queries, and chatbot operational interfaces
– Monitor LLMOps pipelines — tracking prompt execution logs, model API health, token usage, and cost anomalies — escalating degradation or unexpected behaviour promptly
– Perform root cause analysis (RCA) for recurring platform incidents across observability, IaC, CI/CD, ITSM, and AI tooling — documenting findings and driving preventive actions
– Support audit, regulatory, and compliance activities — producing platform access logs, configuration exports, and change evidence for PCI DSS, SOX, and FSI regulatory review
Professional & Technical Skills:
Must-Have Technical Skills
Observability Operations: Splunk — dashboard management, alert rule configuration, metric scraper troubleshooting, and log pipeline operations (ELK/Splunk)
Terraform Operations: Plan/apply execution, state file management, drift alert response, module consumption support, and workspace administration across AWS and GCP environments
CI/CD Operations: GitHub Actions and ArgoCD — pipeline health monitoring, runner management, deployment failure triage, and GitOps workflow support for infrastructure and application teams
HashiCorp Vault: Secret lease monitoring, auth method health, cluster status monitoring,
and escalation-ready diagnosis of secrets management failures
Xmatters: On-call schedule management, escalation policy configuration, alert routing, and integration health with upstream monitoring platforms
AI Tooling Operations: LLM runbook pipeline monitoring, agentic ITSM workflow health, LLMOps dashboard operation, token/cost anomaly alerting, and AI tool incident triage
Scripting & Automation: Python or Bash — operational scripting for platform task automation, API-driven tooling management, and alert-driven remediation workflows
Incident & Change Management: Structured P1–P3 platform incident handling, ITIL-aligned change execution, post-change verification, and RCA documentation in FSI environments
CMDB & Asset Management: Automated discovery tool operations, asset record accuracy monitoring, lifecycle tracking, and CMDB reconciliation for infrastructure tooling assets
Preferred / Advantageous
– Familiarity with LLM APIs (OpenAI, Anthropic, Google Gemini)
— sufficient to triage AI tooling incidents and support engineering teams with prompt execution issues
– Experience with Ansible or Chef for configuration management task execution and playbook troubleshooting – Exposure to Kubernetes operations (EKS or GKE)
— sufficient to support platform tooling deployed on container platforms
– Familiarity with observability-as-code practices — dashboard and alert rules managed as version-controlled configurations
Additional Information:
– Observability, IaC, CI/CD, ITSM, and AI tooling platforms are available and performing — incidents are resolved within SLA with accurate RCA and preventive actions that reduce recurrence
– Engineering teams across Cloud, Network, Security, Database, and Voice towers experience reliable, responsive platform tooling — support requests are resolved promptly and platform outages are communicated clearly
– Changes across the tooling estate are executed cleanly — validated, implemented without unplanned impact, and verified post-completion with documented evidence
– AI-augmented operations tooling is operational and monitored — LLM pipelines, agentic workflows, and alert correlation services are healthy, and degradation is escalated before it impacts engineering teams
– Platform access, audit logs, and CMDB records are consistently accurate — supporting PCI DSS, SOX, and regulatory review without remediation effort
- The candidate should have minimum 3 years of experience in Agentic Orchestration.
- This position is based at our Bengaluru office.
- A 15 years full time education is required.
15 years full time education
📌 Operations Engineer (India)
🏢 Accenture
📍 India