29 Aug
|
Talent Socio
|
India
29 Aug
Talent Socio
India
About the Role:
As part of our client's AIOps program, AIOps Support Engineers perform hands-on, manual triage of alarms and monitoring alerts across a large and diverse application portfolio. The role starts with 50 applications and grows to roughly 320 over two years, spanning different technology stacks, different log aggregation tools, and different lifecycle dispositions (Invest, Tolerate, Retire, Migrate). The focus today is disciplined, high-quality manual triage with a clear roadmap toward more automated, proactive detection as the underlying telemetry backbone (built on Open Telemetry) matures. This is not an Agentic-automation role; it is the foundational human layer that makes accurate, standardized data and eventually automation possible.
What You'll Do
Perform manual Tier 1 triage: monitor, detect, classify, and route incidents based on alarm and alert signals from a range of source applications.
Work across a diverse toolset: interpret alerts from multiple log aggregation and observability platforms Dynatrace, Current Relic, Manage Engine, Glass box, and others and translate them into consistent, actionable incident data.
Support standardized, real-time telemetry help build and maintain a clean Open Telemetry-based data backbone, ensuring alarm and alert data is standardized, timely, and reliable.
Move from reactive to proactive: contribute to the shift from reactive incident response toward proactive issue prevention by flagging patterns, gaps, and data-quality issues in monitoring coverage.
Understand a complex app ecosystem: learn and track the technology stack, architecture disposition (Invest, Tolerate, Retire, Migrate), and ownership of each supported application.
Engage support groups and app owners:
coordinate with the correct support group and application owner for each application to resolve or escalate incidents quickly.
Support a growing, moving target: adapt as new applications are onboarded the portfolio will grow from 50 to approximately 320 applications over two years.
Practice data stewardship: flag inconsistent, noisy, or low-quality alert data and help drive it toward a clean, standardized state. Key Technical Competencies
Application & Production Support (L1/L1.5)
Incident, Problem & Change Management (ITIL)
Application, Middleware & Database Log Analysis
Monitoring & Observability (Dynatrace, New Relic, AWS Cloud Watch, and similar tools)
Open Telemetry & Distributed Tracing (working knowledge)
Proactive Alert Monitoring & Outage Prevention
Linux & Windows OS Troubleshooting
Basic Networking (TCP/IP, DNS, HTTP/HTTPS, SSL, Load Balancers)
SQL & Database Query Analysis
API & Integration Troubleshooting
Hybrid Environment Support (On-Premises + Cloud)
Service Now / Jira Ticket Management
What We're Looking For
Experience: 2 5 years in application or production support (L1/L1.5), preferably supporting hybrid on-premises and cloud environments.
Willingness to do manual, hands-on triage: comfortable with the detail work of alert monitoring and classification, especially in the early stages of the program before automation scales.
Adaptability: able to work across a complex, changing application ecosystem different technology stacks, different log aggregation tools, and different architecture placements.
Customer-centric mindset: focused on service stability, operational excellence, and continuous improvement.
Clear, fast communication: able to interpret alert data with speed and clarity and turn it into actionable next steps for support groups and application owners.
Growth orientation: interested in developing toward more proactive, automation-supported monitoring as the program matures.
Must-Have Skills
Understanding of the Incident Management Lifecycle, and Problem, Change Request, and Service Request concepts (ITIL)
Basic CMDB concepts
Hands-on experience working with log aggregation technologies (Dynatrace, New Relic, Manage Engine, Glassbox, or similar)
Working knowledge of JSON and XML, and basic file/task automation
Knowledge of IT infrastructure and basic networking: VMs, firewalls, load balancers, containers, Open Shift (OCP), Kubernetes
Unix Shell scripting and Windows batch file creation
Basic knowledge of cloud concepts, storage, and security
Basic understanding of TLS, SSL, tokens, and secret management
Familiarity with API gateways and API toolkits (Postman, SOAP UI, or similar)
Ability to support outage response and work effectively with cross-functional teams
Ability to contribute to Root Cause Analyses (RCAs) and knowledge/runbook creation
Basic understanding of SLA/SLO concepts, reporting, and availability calculation
Basic understanding of data concepts: data latency, data fragmentation, data lineage, and data marts
📌 AIOps Support Engineer (India)
🏢 Talent Socio
📍 India