20 Aug
|
Accenture
|
Bengaluru
20 Aug
Accenture
Bengaluru
Project Role : Operations Engineer
Project Role Description : Support the operations and/or manage delivery for production systems and services based on operational requirements and service agreement.
Must have skills : Agentic Orchestration, Agentic AI, AI Orchestration LLM
Positive to have skills : NA
Minimum 3 year(s) of experience is required
Educational Qualification : 15 years full time education
Summary:
A Tools Platforms Service Engineer is responsible for the day-to-day operation, monitoring, support, and maintenance of the infrastructure engineering tooling estate - spanning observability platforms, infrastructure-as-code tooling, CI/CD pipelines, ITSM platforms, internal developer portals, secret management, and AI-augmented operations tooling. The role ensures that the platforms enabling Cloud, Network, Security, Database, and Voice teams to operate at scale remain available, performant, and fit for purpose.
At Level 9 / 10, this individual forms the primary operational execution layer within the tools and platforms tower - a proficient hands-on engineer who keeps platform tooling running reliably, resolves incidents and service requests efficiently, executes changes cleanly, and supports the adoption of AI-augmented operational capabilities across the infrastructure engineering function.
PLATFORMS IN SCOPE
Observability:
- ELK/Splunk
- OpenTelemetry
- Alertmanager
Automation IaC:
- Terraform (consumer)
- Ansible playbooks
- GitHub Actions
- HashiCorp Vault
ITSM DevOps:
- ServiceNow
- Jira/Confluence
- Xmatters
- CMDB tooling
AI-Augmented Ops:
- LLM runbook support
- AI alert triage
- Chatbot ops interfaces
- Prompt execution
- LLMOps monitoring
Roles Responsibilities:
Observability Monitoring Platform Operations
- Monitor and maintain observability platforms - Splunk/ELK - ensuring dashboards, alert rules, and data pipelines remain operational and accurate across all infrastructure tiers
- Triage and resolve observability platform incidents - broken scrapers, missing metrics, failed log ingestion, dashboard errors - within defined SLA commitments
- Manage alert configuration - creating, updating, and tuning alert rules in Alertmanager, or equivalent to reduce noise and maintain actionable alerting across infrastructure teams
- Operate ELK stack or Splunk log pipelines - managing index lifecycle, ingest pipeline health, and log source connectivity for infrastructure platforms
- Execute observability platform changes - dashboard deployments, alert rule updates,
retention policy changes - following approved change management procedures
- Maintain runbooks and operational documentation for all observability platforms - keeping procedures current, accurate, and audit-ready
ITSM, DevOps Tooling AI-Augmented Operations
- Manage Xmatters - maintaining escalation policies, on-call schedules, alert routing rules, and integration health with upstream observability platforms
- Support Jira and Confluence platform operations - managing project configurations, workflow rules, automation triggers, and user access across engineering teams
- Monitor and support AI-augmented operations tooling - LLM runbook automation pipelines, agentic ITSM workflow health, and AI alert correlation services - ensuring AI tools are operational and escalating failures to Principal Engineers
- Support engineering teams using AI-augmented operations capabilities - triaging issues with LLM-powered runbook execution, RAG knowledge base queries, and chatbot operational interfaces
- Monitor LLMOps pipelines - tracking prompt execution logs, model API health, token usage, and cost anomalies - escalating degradation or unexpected behaviour promptly
- Perform root cause analysis (RCA) for recurring platform incidents across observability, IaC, CI/CD, ITSM, and AI tooling - documenting findings and driving preventive actions
- Support audit, regulatory, and compliance activities - producing platform access logs, configuration exports, and change evidence for PCI DSS, SOX, and FSI regulatory review
Professional Technical Skills:
Must-Have Technical Skills
Observability Operations: Splunk - dashboard management, alert rule configuration, metric scraper troubleshooting, and log pipeline operations (ELK/Splunk)
Terraform Operations: Plan/apply execution, state file management, drift alert response, module consumption support, and workspace administration across AWS and GCP environments
CI/CD Operations: GitHub Actions and ArgoCD - pipeline health monitoring, runner management, deployment failure triage, and GitOps workflow support for infrastructure and application teams
HashiCorp Vault: Secret lease monitoring, auth method health, cluster status monitoring, and escalation-ready diagnosis of secrets management failures
Xmatters: On-call schedule management, escalation policy configuration, alert routing, and integration health with upstream monitoring platforms
AI Tooling Operations: LLM runbook pipeline monitoring, agentic ITSM workflow health, LLMOps dashboard operation, token/cost anomaly alerting, and AI tool incident triage
Scripting Automation: Python or Bash - operational scripting for platform task automation, API-driven tooling management, and alert-driven remediation workflows
Incident Change Management: Structured P1-P3 platform incident handling, ITIL-aligned change execution, post-change verification, and RCA documentation in FSI environments
CMDB Asset Management: Automated discovery tool operations, asset record accuracy monitoring, lifecycle tracking, and CMDB reconciliation for infrastructure tooling assets
Preferred / Advantageous
- Familiarity with LLM APIs (OpenAI, Anthropic, Google Gemini)
- sufficient to triage AI tooling incidents and support engineering teams with prompt execution issues
- Experience with Ansible or Chef for configuration management task execution and playbook troubleshooting - Exposure to Kubernetes operations (EKS or GKE)
- sufficient to support platform tooling deployed on container platforms
- Familiarity with observability-as-code practices - dashboard and alert rules managed as version-controlled configurations
- Observability, IaC, CI/CD, ITSM, and AI tooling platforms are available and performing - incidents are resolved within SLA with accurate RCA and preventive actions that reduce recurrence
- Engineering teams across Cloud, Network, Security, Database, and Voice towers experience reliable, responsive platform tooling - support requests are resolved promptly and platform outages are communicated clearly
- Changes across the tooling estate are executed cleanly - validated, implemented without unplanned impact, and verified post-completion with documented evidence
- AI-augmented operations tooling is operational and monitored - LLM pipelines, agentic workflows, and alert correlation services are healthy, and degradation is escalated before it impacts engineering teams
- Platform access, audit logs, and CMDB records are consistently accurate - supporting PCI DSS, SOX, and regulatory review without remediation effort
- The candidate should have minimum 3 years of experience in Agentic Orchestration.
- This position is based at our Bengaluru office.
- A 15 years full time education is required.
Qualification 15 years full time education
📌 Operations Engineer (Bengaluru)
🏢 Accenture
📍 Bengaluru