16 Sep
|
Mattel
|
Hyderabad
Job Description
The Nagios Engineer manages and enhances Nagios systems to monitor enterprise infrastructure, networks, and applications. This role ensures accurate alerting, efficient system performance, and ongoing platform reliability.
Key Responsibilities
- Install, configure, and maintain Nagios Core, Nagios XI monitoring solutions
- Upgrade Nagios and related plugins for system and network monitoring.
- Good hands on experience in Linix/ Unix servers.
- Develop and refine monitoring scripts and templates for consistent coverage.
- Monitor and troubleshoot: Linux, Windows and applications
- Configure and manage: Hosts, services, host groups, and service groups, Active and passive checks, Alerts, escalations, and notifications
- Integrate Nagios XI with: SNMP, NRPE, NCPA, NSClient++, Ticketing tools (ServiceNow, Jira, etc.), Email, SMS, and collaboration tools
- Perform root cause analysis (RCA) and provide incident reports
- Optimize monitoring performance and reduce alert noise (alert fatigue)
- Build dashboards, reports, and capacity planning metrics
- Collaborate with infrastructure, application, and security teams
- Create and maintain monitoring documentation and SOPs
- Support audits,
compliance requirements, and best operational practices
- Analyze performance metrics to identify trends and optimization opportunities.
- Creating dashboards and reports.
- Maintain SOX-compliant documentation for monitoring configurations and changes.
- Collaborate with DC Ops, NOC, and infrastructure teams to resolve system issues.
- Document monitoring standards and ensure consistent best practices.
Qualifications
- Bachelor’s degree in IT or equivalent experience.
- 5 - 8 years of experience managing Nagios XI in enterprise environments.
- Robust Linux and scripting knowledge (Bash, Python, Perl).
- Experience with SNMP, network protocols, and performance metrics.
- Strong analytical and problem-solving abilities.
- Experience with agent-based and agentless monitoring
- Familiarity with performance monitoring and capacity management
- Knowledge of log analysis and alert correlation
Tools & Technologies (Nice to Have)
Other monitoring tools: SCOM, SolarWinds, Prometheus, Grafana, AppDynamics
Cloud monitoring (AWS, Azure, GCP)
ITSM tools (ServiceNow)
📌 Nagios Engineer (Hyderabad)
🏢 Mattel
📍 Hyderabad