19 Aug
|
Accenture in India
|
Bengaluru
19 Aug
Accenture in India
Bengaluru
Project Role : Operations Engineer
Project Role Description : Support the operations and/or manage delivery for production systems and services based on operational requirements and service agreement.
Must have skills : Agentic Orchestration, Agentic AI, AI Orchestration &LLM;
Good to have skills : NA
Minimum 3 Year(s) Of Experience Is Required
Educational Qualification : 15 years full time education
Summary
A Tools &
- Platforms Service Engineer is responsible for the day-to-day operation, monitoring, support, and maintenance of the infrastructure engineering tooling estate — spanning observability platforms, infrastructure-as-code tooling, CI/CD pipelines, ITSM platforms, internal developer portals, secret management, and AI-augmented operations tooling. The role ensures that the platforms enabling Cloud, Network, Security, Database, and Voice teams to operate at scale remain available, performant, and fit for purpose.
At Level 9 / 10, this individual forms the primary operational execution layer within the tools and platforms tower — a proficient hands-on engineer who keeps platform tooling running reliably, resolves incidents and service requests efficiently, executes changes cleanly, and supports the adoption of AI-augmented operational capabilities across the infrastructure engineering function.
PLATFORMS IN SCOPE
Observability
–
- ELK/Splunk
–
- OpenTelemetry
–
- Alertmanager
Automation &
- IaC:
–
- Terraform (consumer)
–
- Ansible playbooks
–
- GitHub Actions
–
- HashiCorp Vault
ITSM &
- DevOps:
–
- ServiceNow
–
- Jira/Confluence
–
- Xmatters
–
- CMDB tooling
AI-Augmented Ops:
–
- LLM runbook support
–
- AI alert triage
–
- Chatbot ops interfaces
–
- Prompt execution
–
- LLMOps monitoring
Roles &
- Responsibilities:
Observability &
- Monitoring Platform Operations
–
- Monitor and maintain observability platforms —
- Splunk/ELK — ensuring dashboards, alert rules, and data pipelines remain operational and accurate across all infrastructure tiers
–
- Triage and resolve observability platform incidents — broken scrapers, missing metrics, failed log ingestion, dashboard errors — within defined SLA commitments
–
- Manage alert configuration — creating, updating, and tuning alert rules in Alertmanager, or equivalent to reduce noise and maintain actionable alerting across infrastructure teams
–
- Operate ELK stack or Splunk log pipelines — managing index lifecycle, ingest pipeline health, and log source connectivity for infrastructure platforms
–
- Execute observability platform changes — dashboard deployments, alert rule updates,
retention policy changes — following approved change management procedures
–
- Maintain runbooks and operational documentation for all observability platforms — keeping procedures current, accurate, and audit-ready
ITSM, DevOps Tooling &
- AI-Augmented Operations
–
- Manage Xmatters — maintaining escalation policies, on-call schedules, alert routing rules, and integration health with upstream observability platforms
–
- Support Jira and Confluence platform operations — managing project configurations, workflow rules, automation triggers, and user access across engineering teams
–
- Monitor and support AI-augmented operations tooling —
- LLM runbook automation pipelines, agentic ITSM workflow health, and AI alert correlation services — ensuring AI tools are operational and escalating failures to Principal Engineers
–
- Support engineering teams using AI-augmented operations capabilities — triaging issues with LLM-powered runbook execution, RAG knowledge base queries, and chatbot operational interfaces
–
- Monitor LLMOps pipelines — tracking prompt execution logs, model API health, token usage, and cost anomalies — escalating degradation or unexpected behaviour promptly
–
- Perform root cause analysis (RCA) for recurring platform incidents across observability, IaC, CI/CD, ITSM, and AI tooling — documenting findings and driving preventive actions
–
- Support audit, regulatory, and compliance activities — producing platform access logs, configuration exports, and change evidence for PCI DSS, SOX, and FSI regulatory review
Qualified &
- Technical Skills:
Must-Have Technical Skills
Observability Operations: Splunk — dashboard management, alert rule configuration, metric scraper troubleshooting, and log pipeline operations (ELK/Splunk)
Terraform Operations: Plan/apply execution, state file management, drift alert response, module consumption support, and workspace administration across AWS and GCP environments
CI/CD Operations: GitHub Actions and ArgoCD — pipeline health monitoring, runner management, deployment failure triage, and GitOps workflow support for infrastructure and application teams
HashiCorp Vault: Secret lease monitoring, auth method health, cluster status monitoring,
and escalation-ready diagnosis of secrets management failures
Xmatters: On-call schedule management, escalation policy configuration, alert routing, and integration health with upstream monitoring platforms
AI Tooling Operations: LLM runbook pipeline monitoring, agentic ITSM workflow health, LLMOps dashboard operation, token/cost anomaly alerting, and AI tool incident triage
Scripting &
- Automation: Python or Bash — operational scripting for platform task automation, API-driven tooling management, and alert-driven remediation workflows
Incident &
- Change Management: Structured P1–P3 platform incident handling, ITIL-aligned change execution, post-change verification, and RCA documentation in FSI environments
CMDB &
- Asset Management: Automated discovery tool operations, asset record accuracy monitoring, lifecycle tracking, and CMDB reconciliation for infrastructure tooling assets
Preferred / Advantageous
–
- Familiarity with LLM APIs (OpenAI, Anthropic, Google Gemini)
— sufficient to triage AI tooling incidents and support engineering teams with prompt execution issues
–
Experience with Ansible or Chef for configuration management task execution and playbook troubleshooting –
- Exposure to Kubernetes operations (EKS or GKE)
— sufficient to support platform tooling deployed on container platforms
–
- Familiarity with observability-as-code practices — dashboard and alert rules managed as version-controlled configurations
Additional Information
–
- Observability, IaC, CI/CD, ITSM, and AI tooling platforms are available and performing — incidents are resolved within SLA with accurate RCA and preventive actions that reduce recurrence
–
- Engineering teams across Cloud, Network, Security, Database, and Voice towers experience reliable, responsive platform tooling — support requests are resolved promptly and platform outages are communicated clearly
–
- Changes across the tooling estate are executed cleanly — validated, implemented without unplanned impact, and verified post-completion with documented evidence
–
- AI-augmented operations tooling is operational and monitored —
- LLM pipelines, agentic workflows, and alert correlation services are healthy, and degradation is escalated before it impacts engineering teams
–
- Platform access, audit logs, and CMDB records are consistently accurate — supporting PCI DSS, SOX, and regulatory review without remediation effort
- The candidate should have minimum 3 years of experience in Agentic Orchestration.
- This position is based at our Bengaluru office.
- A 15 years full time education is required.
📌 Operations Engineer (Bengaluru)
🏢 Accenture in India
📍 Bengaluru