31 Jul
|
iKrux
|
Bengaluru
Full Stack Observability Engineer
Experience: 69 Years
Location: Bangalore / Hyderabad / Chennai / Pune (As per Business Requirement)
Job Summary
We are looking for a highly skilled Full Stack Observability Engineer with hands-on experience in enterprise observability platforms, especially Datadog and Dynatrace. The ideal candidate should have expertise in monitoring infrastructure, applications, networks, and cloud environments while building intelligent dashboards, alerting mechanisms, and automation solutions. Experience with Python/PowerShell scripting and AIOps concepts is highly preferred.
Key Responsibilities
- Design, implement, and maintain end-to-end observability solutions across infrastructure, applications, networks, and cloud environments.
- Administer and optimize observability platforms including Datadog and Dynatrace.
- Configure dashboards, alerts, metrics, logs, distributed tracing, and performance monitoring.
- Develop automation scripts using Python and PowerShell to improve monitoring efficiency.
- Perform root cause analysis (RCA) and troubleshoot application and infrastructure performance issues.
- Implement intelligent alerting strategies to minimize alert fatigue and improve incident response.
- Collaborate with development, infrastructure, and operations teams to enhance platform reliability and performance.
- Monitor system health, capacity, and availability using industry best practices.
- Support predictive monitoring and AIOps initiatives for proactive issue detection.
- Create technical documentation and operational runbooks.
Required Skills
- 6–9 years of experience in Observability, Monitoring, or Site Reliability Engineering.
- Strong hands-on experience with Datadog and Dynatrace (Mandatory).
- Experience with monitoring tools such as Prometheus, Grafana, and New Relic.
- Strong knowledge of Infrastructure Monitoring, Application Performance Monitoring (APM), and Distributed Tracing.
- Proficiency in Python and PowerShell scripting.
- Experience in Log Management, Metrics Collection, Alerting, and Dashboard Development.
- Positive understanding of Linux, Windows, networking, and cloud infrastructure.
- Strong analytical, troubleshooting, and Root Cause Analysis (RCA) skills.
- Excellent communication and stakeholder management skills.
Preferred Skills
- Experience with AIOps, anomaly detection, and event correlation.
- Exposure to predictive monitoring and capacity planning.
- Knowledge of cloud platforms such as AWS, Azure, or GCP.
- Experience working in enterprise production environments.
Mandatory Skills
- Datadog
- Dynatrace
- Python
- PowerShell
- Observability
- Infrastructure Monitoring
- Application Monitoring
- APM
- Prometheus
- Grafana
Good to Have
- New Relic
- AIOps
- Distributed Tracing
- Predictive Maintenance
- Event Correlation
- Capacity Planning
- Cloud Monitoring (AWS/Azure/GCP)
Notice Period : Immediate Keywords: Datadog, Dynatrace, Observability, Monitoring, APM, Site Reliability Engineer, SRE, Prometheus, Grafana, New Relic, Python, PowerShell, Infrastructure Monitoring, Application Monitoring, Distributed Tracing, Dashboards, Alerting, Root Cause Analysis, AIOps.
📌 Full Stack Observability Engineer (Bengaluru)
🏢 iKrux
📍 Bengaluru