About the Role:
OEC is in the middle of a multi-year journey to modernise its infrastructure and applications, transforming how we build, deploy, and operate software at global scale. The Operations Specialist (L2) is a core member of the Monitoring Team within the Enterprise Operations Team, responsible for the configuration, maintenance, and continuous improvement of OEC’s observability stack, primarily powered by Datadog. In this role you will own L2 alert triage, incident response, dashboard development, and service onboarding — contributing to the team’s tooling standards and best practices. You will also handle day-to-day server monitoring, infrastructure support, and service management tasks that keep OEC’s global platforms running reliably.
Key Responsibilities & Duties (essential to the job)
Monitoring & alerting
- Design, configure, and maintain Datadog monitors, composite alerts, and notification channels for production and non-production environments.
- Own the L2 alert triage process — investigate, diagnose, and resolve escalated alerts from L1, ensuring timely root cause identification.
- Tune alert thresholds and suppression rules to minimise noise and reduce false positive rates.
- Manage SLO (Service Level Objective) definitions, tracking, and reporting across assigned services.
- Participate in the on-call rotation for critical monitoring alerts and P1/P2 incident response.
Dashboards & observability
- Build and maintain Datadog dashboards covering infrastructure health, application performance, log analytics, and business KPIs.
- Configure and manage Datadog APM (Application Performance Monitoring), distributed tracing, and error tracking for key services.
- Implement and manage log ingestion pipelines, parsing rules, and log-based monitors.
- Develop and maintain Synthetic monitors (API and browser tests) for uptime and user experience validation.
Service onboarding & integrations
- Onboard new services and infrastructure components into the Datadog monitoring framework in collaboration with DevOps, Cloud, and Engineering teams.
- Configure and maintain Datadog integrations with cloud platforms (AWS, Azure, Rackspace), Kubernetes, containerised workloads, and third-party tools.
- Support the migration of legacy monitoring tooling (SCOM, Dynatrace, Grafana, Prometheus, Pingdom, Apigee, Redgate, IDERA) into Datadog.
Incident response & runbooks
- Act as L2 escalation point during incidents — perform deep-dive investigations using metrics, logs, traces, and dashboards.
- Create, maintain, and improve runbooks for common alert scenarios, incident response procedures, and post-incident remediation steps.
- Contribute to post-incident reviews, documenting root cause findings and identifying monitoring gaps to prevent recurrence.
- Alerts system owners,
stakeholders, and management of degraded system status and Priority 1 and Priority 2 incidents; issues tickets for incident and problem escalation.
AI & automation
- Leverage Datadog AI capabilities including Watchdog, anomaly detection, outlier detection, and Bits AI to improve proactive issue detection.
- Implement forecast-based monitors for capacity management (CPU, memory, disk).
- Use AI-based alert correlation and deduplication features to reduce alert fatigue.
- Contribute to automation of repetitive monitoring tasks using Datadog’s API and scripting tools.
Standards & governance
- Adhere to and actively contribute to monitoring standards, tagging strategies, and naming conventions.
- Participate in regular alert quality reviews, dashboard audits, and SLO compliance checks.
- Maintain accurate documentation of monitoring configurations, integrations, and architectural decisions.
- Creates and maintains knowledge articles to be used internally and/or externally for training, best practices, solutions, or processes relating to applications, environments, and related technologies.
Infrastructure operations & support
- Analyses and troubleshoots Microsoft and UNIX/Linux server configurations and processes.
- Performs moderately complex database administration.
- Monitors data centre networks, infrastructure bandwidth, application health, servers, and other infrastructure; coordinates and communicates with vendors.
- Diagnoses and researches (using knowledge base) application incidents, monitoring alerts, and service requests; provides assistance and guidance to associate operations specialists.
- Adheres to all incident and service request processes and procedures in accordance with established Service Level Agreements (SLAs).
- Diagnoses, researches, and resolves Level-2 technical hardware and software incidents, monitoring alerts, and service requests.
- Works on ad-hoc projects to support Infrastructure Engineering or other departments, as requested.
- Demonstrates a adaptable and adaptable approach to work and adjusts to shifts in priorities as the needs of the business change.
- Collaborates with DevOps, Cloud, and application teams to align monitoring coverage with service requirements.
Education
A bachelor’s degree from an accredited college or university in Computer Science, Information Technology, Engineering, or a related discipline is required.
In the absence of a degree, equivalent work experience directly related to the key responsibilities of the role will be considered as a substitute for the degree.
Experience, Skills and Key Competencies
At least 3–5 years of experience in an infrastructure, DevOps, SRE, or monitoring engineering role is required, with hands-on experience in Datadog and working knowledge of cloud infrastructure and containerised environments. Must also be able to demonstrate the following skills and abilities:
Technical skills
- Hands-on experience with Datadog — monitors, dashboards, APM, log management, SLOs, and synthetic monitoring.
- Solid understanding of cloud infrastructure (AWS required; Azure or GCP a plus) and containerised environments (Docker, Kubernetes).
- Familiarity with observability concepts: metrics, traces, and logs.
- Experience with at least one scripting language (Python, Bash, or PowerShell) for automation and tooling.
- Working knowledge of monitoring and alerting tools such as Prometheus, Grafana, Dynatrace, or similar platforms.
- Understanding of ITSM concepts and ticketing workflows (Jira, JSM, or ServiceNow).
- Good working knowledge of Microsoft and UNIX/Linux server configurations, Active Directory, and internet protocols (HTTP, FTP, DNS, IP, SSH).
Soft skills & competencies
- Strong analytical and problem-solving skills with the ability to diagnose complex infrastructure issues under pressure.
- Clear written and verbal communication skills; able to produce runbooks and incident summaries for both technical and non-technical audiences.
- Collaborative team player who can work effectively across Engineering, DevOps, and Operations disciplines.
- Self-motivated with a continuous improvement mindset and a proactive approach to identifying and resolving issues before they escalate.
- Adaptable and adaptable approach to work; able to adjust to shifts in priorities as the needs of the business change.
Preferred Qualifications
- Datadog Fundamentals or Datadog APM certification (or actively working towards one).
- At least 1 year of experience with CI/CD pipelines (Jenkins, GitHub Actions, or Azure DevOps).
- Knowledge of network monitoring and distributed systems architecture.
- Exposure to ITIL or ITSM frameworks (Incident, Problem, and Change Management).
- Familiarity with Atlassian tools (Jira, JSM, Confluence) for documentation and ticket management.
Special Position Requirements
- Must be able to read, write, understand, and speak fluent English.
- Must be able to work flexible shifts including holidays and weekends, to provide support across locations and time zones.
- Must be able to participate in an on-call rotation for critical monitoring alerts and P1/P2 incident response.
📌 Operations Specialist (Chennai)
🏢 OEC
📍 Chennai