07 Aug
|
Analog Devices
|
Bengaluru
07 Aug
Analog Devices
Bengaluru
Key Responsibilities
- Design, implement, and maintain comprehensive monitoring strategies across Linux, Storage, Applications, and Networking platforms
- Define and establish KPI/Metrics standards and best practices for the organization, ensuring consistency and measurability
- Architect and deploy monitoring solutions using industry-leading tools including Prometheus, Grafana, Nagios, and SolarWinds
- Develop and implement Kafka-based event streaming and monitoring pipelines for real-time data collection and analysis
- Integrate and manage AI monitoring capabilities to enable intelligent alerting, anomaly detection, and predictive insights
- Create and maintain comprehensive documentation and runbooks for monitoring infrastructure, troubleshooting procedures, and operational playbooks
- Develop custom scripts and automation in bash, Python, or Go to enhance monitoring capabilities and reduce operational overhead
- Establish and refine SLAs, SLOs, and alerting thresholds based on business requirements and operational data
- Conduct performance analysis and optimization of monitoring infrastructure to reduce latency and improve system reliability
- Collaborate with DevOps, Platform, and Application teams to ensure comprehensive observability across the entire stack
- Stay current with emerging monitoring and observability trends, tools, and technologies
- Drive continuous improvement initiatives by analyzing metrics, gathering feedback, and implementing enhancements to the user experience
Required Qualifications
- Minimum 8+ years of professional experience in monitoring, observability, systems engineering, or DevOps roles
- Proven experience leading and mentoring technical teams in monitoring or infrastructure domains
- Deep hands-on expertise with opensource monitoring tools such as Prometheus, Grafana, and Nagios
- Strong experience with enterprise monitoring platforms like SolarWinds
- Proficiency in bash scripting and Linux system administration (RHEL/CentOS/Ubuntu preferred)
- Experience with event-driven architectures and Apache Kafka for real-time data processing
- Knowledge of monitoring and observability across Linux systems, storage infrastructure, applications, and network platforms
- Understanding of metrics collection, time-series data management, and data-driven decision-making
- Experience implementing AI/ML-based monitoring solutions or anomaly detection systems
- Solid problem-solving skills and ability to work effectively in fast-paced, complex environments
- Excellent communication skills with ability to present technical concepts to both technical and non-technical stakeholders
- Experience with Infrastructure-as-Code (IaC) tools (Terraform, Ansible, etc.)
Preferred Qualifications
- Experience with containerized environments and Kubernetes monitoring
- Proficiency in Python or Go for custom tooling development
- Experience with cloud platforms (AWS, Azure, GCP) and their native monitoring services
- Experience with observability backends such as OpenTelemetry.
- Knowledge of application performance monitoring (APM) tools
- Experience with incident response and post-mortem processes
- Certification in relevant areas (e.g., Kubernetes, cloud platforms, monitoring tools)
- Contribution to open-source monitoring projects
Required Technical Skills
- Monitoring Platforms: Prometheus, Grafana,
Nagios, Azure Monitor, AWS CloudWatch, SolarWinds
- Data Streaming: Apache Kafka, event processing pipelines
- AI Monitoring: Anomaly detection, intelligent alerting, ML-based insights
- Scripting Languages: Bash, Python, Go
- Operating Systems: Linux (RHEL, CentOS, Ubuntu), system administration
- Infrastructure: Storage systems, networking, application monitoring
- Data Visualization: Grafana dashboards, custom visualizations
- Cloud & Containerization: Docker, Kubernetes (optional but valuable)
Strategic Responsibilities
As a leader in this role, you will be responsible for:
- Developing and communicating the monitoring and observability strategy aligned with organizational goals
• Setting industry-standard KPIs and metrics that drive business value and operational excellence
• Using data-driven approaches to identify improvement opportunities and measure success
• Building a scalable, maintainable monitoring infrastructure that grows with the organization
• Establishing clear performance baselines and continuous improvement metrics
• Ensuring that monitoring initiatives directly improve user experience and system reliability
What We're Looking For
We seek a hands-on technical leader who is passionate about observability and excellence. Ideal candidates will:
- Balance strategic thinking with hands-on technical expertise
• Demonstrate a track record of building high-performing teams
• Show deep curiosity about system behavior and data-driven insights
• Possess strong written and verbal communication skills
• Take pride in mentoring others and sharing knowledge
• Approach challenges with creativity and persistence
• Stay current with evolving monitoring and observability technologies
📌 Sr. Monitoring and Observability Engineer (Bengaluru)
🏢 Analog Devices
📍 Bengaluru