Lead the implementation of enterprise observability for applications, APIs, services, batch jobs, and data pipelines. Design and standardize monitoring, alerting, logging, metrics, and health checks across distributed systems. Integrate observability platforms with incident management and automation tools to support proactive issue detection and remediation. Support reliability and availability of integration platforms built on AWS/Azure Perform advanced troubleshooting using logs, metrics, and traces to resolve production issues. Mentor engineers on observability best practices and platform usage. Collaborate with product, support, and operations teams to improve service stability and delivery. 5+ years of relevant experience in Observability / Monitoring / Reliability Engineering IBM Instana, Dynatrace, AppDynamics, Prometheus, Grafana Monitoring and alerting design Log management and analysis Metrics and distributed tracing Health checks and SLO/SLI concepts Automation and ITSM integration (ServiceNow workflows, incident automation) CI/CD and release management exposure Cloud integration and messaging exposure Automation and ITSM integration (ServiceNow workflows, incident automation)