Job Description : Observability SME SRE Datadog Azure Automation
Role Overview
We are looking for an experienced Observability SME Site Reliability Engineer who will drive endtoend observability strategy monitoring automation and reliability engineering across a hybrid environment Azure OnPrem The ideal candidate must have deep expertise in Datadog administration and configuration handson experience in Azure services and strong automation skills using Terraform GitHub Actions and PowerShell
This role will be responsible for designing observability standards implementing automated monitoring deployments ensuring platform reliability and enabling proactive incident detection remediation
Key Responsibilities
Observability Datadog Primary Focus
Act as the Subject Matter Expert SME for Datadog owning platform administration configurations access control and best practices
Design implement and maintain observability standards across applications infrastructure APIs and data platforms
Automate Datadog agent deployment configurations dashboards alerts synthetics log pipelines and APM instrumentation using
Terraform
GitHub Actions
Set up and optimize monitoring for Azure services including but not limited to
AKS Azure Kubernetes Service
App Services
Function Apps
Virtual Machines
API Management
Databricks
Azure SQL
Define SLOs SLIs error budgets logging standards and distributed tracing patterns
Site Reliability Engineering
Drive operational excellence through automation performance tuning and reliability improvements
Build automated workflows for incident diagnosis remediation and recoveryinitially in Azure using PowerShell
Implement proactive monitoring to reduce MTTR and prevent recurrence of incidents
Partner with application platform and cloud engineering teams to enable observability and SRE best practices
Lead root cause analysis RCA create improvement plans and ensure closure of action items
Automation Infrastructure as Code
Develop modular reusable Terraform modules for Datadog monitoring Azure resources and logging pipelines
Implement continuous deployment pipelines using GitHub Actions for observability components
Maintain version control practices branching strategies and peer review processes
Technical Skills Required
Mandatory
Datadog SMElevel experience Administration agent configuration integration setup dashboards monitors alerts synthetics log pipelines APM RUM
Azure Cloud experience specifically
AKS
Azure App Services
Function Apps
Virtual Machines
API Management
Databricks
Azure SQL Server
Automation IaC
Robust experience with Terraform
GitHub Actions CICD pipelines
PowerShell scripting for automation
Experience with hybrid environments Azure On Prem
Good to Have
Experience with YAML pipelines ARMBicep optional
Understanding of networking load balancing DNS certificates
Experience with SRE concepts like chaos engineering reliability patterns
Experience Required
1012 years of total experience in Cloud DevOps SRE or Observability roles
Minimum 3 years of handson Datadog experience
Minimum 2 years working with Azure cloud workloads
📌 Observability SME (Hyderabad)
🏢 Alike Thoughts
📍 Hyderabad
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.