06 Aug
|
Tidyhire
|
India
Key Responsibilities:
- Lead the design, implementation, and governance of enterprise observability and monitoring solutions.
- Build and maintain monitoring platforms using Prometheus, Grafana, Azure Monitor, Application Insights, Splunk, and OpenTelemetry.
- Develop dashboards, alerts, and SLO/SLI metrics to proactively monitor application and infrastructure health.
- Monitor and optimize the performance, availability, and reliability of cloud platforms, Kubernetes clusters, and business-critical applications.
- Lead major incident response, root cause analysis (RCA), and continuous service improvement initiatives.
- Implement centralized logging, distributed tracing, and end-to-end observability across cloud-native environments.
- Automate monitoring deployment, alerting, and operational workflows using Terraform, Ansible, Python, PowerShell, or Bash.
- Collaborate with Platform Engineering, DevOps, SRE, Security, and Application teams to establish monitoring standards and best practices.
- Drive capacity planning, performance tuning, and cost optimization for monitoring platforms.
- Mentor engineers on observability practices and foster a culture of reliability, automation, and operational excellence.
Key Skills and Experience:
- 1216 years of overall IT experience, with at least 6+ years of hands-on experience in Observability, Monitoring, Site Reliability Engineering (SRE), or Platform Engineering.
- Proven experience in designing, implementing,
and managing enterprise observability and monitoring solutions for large-scale production environments.
- Robust hands-on expertise with Prometheus, Grafana, Azure Monitor, Application Insights, OpenTelemetry, and centralized logging platforms such as Splunk or ELK.
- Extensive experience monitoring Azure cloud infrastructure, Kubernetes (AKS), containerized applications, distributed systems, and cloud-native platforms.
- Hands-on experience in building dashboards, alerts, metrics, logs, and distributed tracing to ensure end-to-end visibility and proactive monitoring.
- Strong understanding of Site Reliability Engineering (SRE) principles, including SLIs, SLOs, Error Budgets, Incident Management, Root Cause Analysis (RCA), performance optimization, and service reliability.
- Experience supporting enterprise production environments, managing major incidents, and driving high availability, operational excellence, and continuous service improvement.
- Proficiency in automation and Infrastructure as Code (IaC) using Terraform, Ansible, Python, PowerShell, or Bash.
- Strong troubleshooting, analytical, communication, stakeholder management, and cross-functional collaboration skills.
Preferred Certifications:
- Microsoft Certified: Azure Administrator (AZ-104)
- Microsoft Certified: Azure Solutions Architect (AZ-305)
- Certified Kubernetes Administrator (CKA)
- Grafana Certified Professional
- Prometheus Certified Associate (PCA)
- ITIL Foundation
- SRE Foundation / SRE Practitioner
📌 Observability & Monitoring Lead - DevOps (India)
🏢 Tidyhire
📍 India