03 Sep
|
HCLTech
|
Hyderabad
Hyderabad, Telangana
Job Summary
Role Summary
The Observability & Datadog Engineer is responsible for designing, implementing, and managing enterprise monitoring and observability solutions using Datadog. The role focuses on proactive monitoring, performance management, incident reduction, root cause analysis, and improving service reliability across infrastructure, applications, cloud platforms, and business services. Internal observability frameworks emphasise Metrics, Events, Logs, and Traces (MELT), alert correlation, anomaly detection, dashboards, and SRE-driven operations.
Key Responsibilities
Key Responsibilities
Design and implement end-to-end observability solutions using Datadog.
Configure and maintain monitoring for:
Infrastructure (Servers, VMs, Network, Storage)
Cloud platforms (Azure, AWS, GCP)
Containers and Kubernetes
Applications and Microservices
Databases and Middleware
Build and manage:
Dashboards
Alerts and Monitors
Service Level Indicators (SLIs)
Service Level Objectives (SLOs)
Implement Application Performance Monitoring (APM), Log Management, Distributed Tracing, and Synthetic Monitoring.
Perform root cause analysis using metrics, logs, traces, and events.
Reduce alert noise through correlation, automation, and threshold optimisation.
Partner with Development, DevOps, Cloud,
and SRE teams to improve application reliability.
Automate monitoring deployment using Terraform, APIs, Python, or Shell scripting.
Support incident, problem, and change management processes.
Define observability standards, governance, and best practices.
Skill Requirements
Required Skills
Datadog
Infrastructure Monitoring
APM (Application Performance Monitoring)
Log Management
Distributed Tracing
RUM (Real User Monitoring)
Synthetic Monitoring
Network Performance Monitoring
Dashboard & Alert Configuration
SLO/SLI Management
Datadog API & Integrations
Technical Skills
Linux & Windows Administration
Kubernetes / OpenShift
Docker Containers
Azure / AWS / GCP
CI/CD Tools
REST APIs
Terraform / Infrastructure as Code
Python, PowerShell, Bash
Observability Skills
Metrics, Logs, Events & Traces (MELT)
Monitoring Strategy
Alert Correlation
Anomaly Detection
Capacity & Performance Management
Incident Management
Root Cause Analysis
Other Requirements
#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-
📌 Tower Lead (Support & Operations) (Hyderabad)
🏢 HCLTech
📍 Hyderabad