23 Sep
|
Allscripts(India) LLP, ultimately a subsidiary of Altera Digital Health Inc.,[Altera India]
|
Pune
23 Sep
Allscripts(India) LLP, ultimately a subsidiary of Altera Digital Health Inc.,[Altera India]
Pune
Key Responsibilities
- Ensure 99.99% or greater infrastructure availability through proactive monitoring, maintenance coordination, operational controls, and disciplined incident management.
- Oversee the health of infrastructure and services, including compute, storage, memory, network, cloud, database, application, endpoint, and integration components.
- Drive automation-first operations through self-healing pipelines, automated remediation playbooks, and infrastructure-as-code patterns that reduce manual toil, standardize recovery actions, and improve mean time to restore service.
- Ensure alerts are acknowledged promptly, validated against known conditions, correctly categorized, prioritized, documented, and routed to the appropriate resolver group.
- Review alert thresholds, suppression rules, correlation logic, maintenance windows, and routing policies to improve signal quality and reduce avoidable noise.
- Direct initial troubleshooting using approved runbooks, knowledge articles, dashboards, logs, and diagnostic tools.
- Coordinate rapid escalation of warning and exception events that indicate service degradation, capacity risk, or potential incident.
- Serve as the primary escalation point for operational incidents, lead root cause analysis, and drive corrective and preventive actions through completion.
- Support major incident response by establishing situational awareness, assigning monitoring actions, maintaining an event timeline, and providing accurate technical updates.
- Analyze recurring alerts, resource trends, capacity indicators, service dependencies,
and monitoring gaps; initiate corrective actions with engineering and problem management teams.
- Own operating procedures, runbooks, escalation matrices, contact lists, shift checklists, and knowledge documentation.
- Produce operational summaries and periodic reports covering service health, significant events, response performance, recurring conditions, and improvement actions.
- Maintain clear shift coverage, handoff, attendance, workload, and escalation expectations across the monitoring function.
- Act as the operational bridge among monitoring engineers, incident management, service owners, application teams, infrastructure teams, vendors, and business stakeholders.
- Lead post-event reviews and continual improvement initiatives focused on automation, observability, runbook quality, engineer readiness, and faster restoration.
Required Qualifications
- Seven or more years of experience in IT operations, infrastructure support, network operations, cloud operations, application support, site reliability, or a related field.
- 2 years of experience leading a Network Operations team or in a team lead capacity. Preferred.
- Robust working knowledge of ITIL monitoring and event management, incident management,
problem management, change enablement, configuration management, and service-level management.
- Hands-on experience with enterprise monitoring, observability, alerting, ticketing, paging, log analysis, and dashboard platforms.
- Hands-on experience integrating Grafana, PagerDuty, ThousandEyes, Azure Monitor, Amazon CloudWatch, or comparable tools, within an enterprise observability ecosystem.
- Hands-on experience troubleshooting Windows Server operating systems, services, performance, connectivity, storage, patching, and related infrastructure issues in production environments.
- Strong knowledge of AWS and Microsoft Azure platforms, with hands-on experience troubleshooting cloud-hosted services.
- Ability to interpret infrastructure and application telemetry, identify service impact, prioritize competing conditions, and make timely decisions under pressure.
- Demonstrated experience managing shifts, on-call coverage, escalations, runbooks, operational metrics, and cross-functional response.
- Clear written and verbal communication skills, including the ability to translate technical conditions into concise business-impact updates.
- Strong leadership, coaching, collaboration, organization, and continuous-improvement skills.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Manager Network Operations (Pune)
🏢 Allscripts(India) LLP, ultimately a subsidiary of Altera Digital Health Inc.,[Altera India]
📍 Pune