02 Aug
|
Northern Trust
|
Pune
02 Aug
Northern Trust
Pune
We are seeking a highly skilled Senior Observability Engineer to support and enhance enterprise observability capabilities across infrastructure, applications, cloud, and business services. This is an individual contributor role focused on hands-on engineering, platform operations, monitoring standardization, automation, and continuous improvement across observability platforms including Dynatrace, Microsoft SCOM, ServiceNow ITOM & Event Management, and related monitoring integrations.
Key Responsibilities:
- Implement, maintain, and continuously improve enterprise observability capabilities across Dynatrace, SCOM, ServiceNow ITOM & Event Management, and supporting monitoring tools.
- Configure and support monitoring for infrastructure, applications, services, databases, middleware, cloud, and hybrid environments to ensure end-to-end visibility and operational stability.
- Develop, tune, and optimize monitoring alerts, dashboards, thresholds, synthetic checks, anomaly detection, and service-level views to improve signal quality and reduce false positives.
- Drive alert hygiene activities, including duplicate alert reduction, threshold tuning, suppression logic, event enrichment, and monitoring standardization across technology teams.
- Support integrations between observability platforms and ServiceNow Event Management, ensuring events are enriched, correlated, deduplicated, and routed effectively for incident response.
- Manage and enhance monitoring integrations from tools such as Dynatrace, SCOM, infrastructure monitoring sources, and application monitoring platforms into ServiceNow Event Management.
- Use Dynatrace capabilities such as service flow, distributed tracing, Davis AI, anomaly detection, problem correlation, dashboards, management zones, tags, and alerting profiles to improve root cause analysis and operational insights.
- Support Microsoft SCOM monitoring activities including management pack configuration, alert rule tuning, agent health checks, monitoring coverage, and operational troubleshooting.
- Build automation scripts and reusable solutions using Python, PowerShell, REST APIs, YAML/JSON, CI/CD pipelines, and other automation frameworks to improve onboarding, monitoring configuration, health checks, reporting, and operational efficiency.
- Contribute to observability-as-code practices by supporting standardized, repeatable, and automated monitoring onboarding patterns.
- Partner with application, infrastructure, cloud, and support teams to onboard new applications and services into enterprise monitoring platforms.
- Troubleshoot monitoring gaps, integration failures, agent issues, event flow problems, and alerting defects across observability systems.
- Support operational reporting and KPI tracking related to alert volume, noise reduction, MTTR improvement, event quality, monitoring coverage, and automation adoption.
- Maintain technical documentation, operational runbooks, configuration standards, troubleshooting guides, and onboarding procedures.
- Participate in incident reviews and problem management discussions to identify opportunities for monitoring improvement and proactive detection.
- Apply ITIL practices and event lifecycle management principles to improve incident quality, operational response, and service reliability.
Required Skills and Experience:
- 12+ years of overall IT experience, with strong hands-on experience in observability, monitoring,
event management, infrastructure operations, application support, or platform engineering.
- Strong practical experience with Dynatrace including OneAgent, dashboards, alerts, management zones, synthetic monitoring, service flow, problem detection, tagging, and Davis AI capabilities.
- Hands-on experience with Microsoft SCOM, including alert configuration, management packs, agent monitoring, rule tuning, infrastructure monitoring, and operational troubleshooting.
- Strong working knowledge of ServiceNow Event Management / ITOM, including event ingestion, event rules, alert correlation, deduplication, enrichment, service mapping awareness, and incident integration.
- Experience integrating monitoring tools with ServiceNow or similar ITSM platforms.
- Good understanding of observability concepts including metrics, logs, traces, events, topology, service health, synthetic monitoring, and full-stack monitoring.
- Solid scripting and automation experience using Python, PowerShell, REST APIs, shell scripting, or similar technologies.
- Experience with CI/CD tools, source control, configuration files, automation workflows, and infrastructure-as-code or observability-as-code practices.
- Familiarity with cloud and hybrid environments such as Azure, AWS, VMware, Windows, Linux, databases, middleware, and enterprise infrastructure platforms.
- Understanding of ITIL processes, especially Incident Management, Problem Management, Change Management, Event Management, and operational support models.
- Strong analytical and troubleshooting skills with the ability to identify monitoring gaps, alert quality issues, and event flow problems.
- Ability to work independently as an individual contributor while collaborating effectively with engineering, operations, application, and support teams.
📌 Sr. Lead - Observability Engineering (Pune)
🏢 Northern Trust
📍 Pune