08 Sep
|
Exsete Consulting
|
Noida
08 Sep
Exsete Consulting
Noida
– SRE (Monitoring & Observability)
Experience: 1 Year
Role: Site Reliability Engineer (SRE) – Monitoring & Observability
Job Summary
We are looking for an SRE with around 1 year of experience in Monitoring and Observability to help ensure the availability, reliability, and performance of applications and infrastructure. The candidate should have hands-on knowledge of ELK Stack, Linux, and Grafana, along with basic incident management and troubleshooting skills.
Key Responsibilities
- Monitor production applications, infrastructure, APIs, and services for availability and performance.
- Create and manage Grafana dashboards to monitor system and application metrics.
- Configure and monitor alerts for critical issues, performance degradation, and service failures.
- Use the ELK Stack (Elasticsearch, Logstash, Kibana) for centralized log monitoring, analysis, and troubleshooting.
- Perform Linux server monitoring and troubleshooting, including CPU, memory, disk, processes, services, and network-related issues.
- Analyze logs and metrics to identify application and infrastructure issues.
- Respond to production alerts and incidents within defined SLAs.
- Perform initial troubleshooting, incident analysis, and escalation to the appropriate teams.
- Monitor Elasticsearch health, indices, ingestion pipelines, and data availability.
- Support incident management, RCA (Root Cause Analysis), and preventive actions.
- Identify recurring issues and suggest improvements to monitoring, alerting,
and system reliability.
- Maintain monitoring dashboards, alerts, runbooks, and operational documentation.
- Work with development and infrastructure teams to improve application reliability and observability.
Required Skills
- 1 year of experience in SRE, Production Support, Monitoring, or Observability.
- Good knowledge of Linux/Unix operating systems and basic troubleshooting commands.
- Hands-on experience with ELK Stack – Elasticsearch, Logstash, and Kibana.
- Experience creating and monitoring Grafana dashboards.
- Understanding of monitoring concepts such as metrics, logs, alerts, SLIs, SLOs, and SLAs.
- Basic knowledge of CPU, memory, disk, network, and application performance monitoring.
- Valuable understanding of incident management and production support processes.
- Basic knowledge of shell scripting is an advantage.
- Good analytical and problem-solving skills.
- Willingness to work in rotational/on-call support, if required.
Good to Have
- Knowledge of Prometheus and alerting.
- Basic understanding of Docker/Kubernetes.
- Familiarity with cloud platforms such as AWS/Azure/GCP.
- Knowledge of CI/CD and automation.
- Experience with ticketing and incident-management tools.
Key Objective
Ensure high availability, reliability, and performance of production systems through proactive monitoring, effective observability, timely incident response, and continuous improvement of monitoring and alerting.
📌 Data Engineer (Noida)
🏢 Exsete Consulting
📍 Noida