Key Areas of Responsibilities
- Own and support monitoring and SRE operations, ensuring system reliability, availability, and performance.
- Build, enhance, and maintain monitoring solutions using ITRS Geneos, Prometheus, Victoria‑Metrics, Elasticsearch, and Grafana.
- Develop, optimize, and maintain alerting rules, dashboards, and observability pipelines.
- Troubleshoot and resolve complex issues during major incidents, providing clear and timely communication.
- Troubleshoot Linux servers (RHEL 7/8/9), including upgrades, configurations, patching, and maintenance, while determining appropriate monitoring requirements for system changes.
- Analyze logs, investigate issues, and perform fault finding to identify performance exceptions.
- Collaborate with engineering, application, and infrastructure teams to improve system resilience, stability, security, efficiency, and scalability.
- Contribute to automation strategies, deployment processes, and continuous operational improvements.
- Participate in on‑call rotations, including off‑hours and scheduled weekend support.
- Participate in Disaster Recovery (DR) and Business Continuity Planning (BCP) drills.
- Continuously research and adopt up-to-date monitoring and SRE tools and practices.
Requirements
- Bachelor’s degree in computer science / engineering
- Minimum 8 years’ experience within IT / Investment bank.
- Strong experience with monitoring and observability platforms, including: ITRS Geneos, Prometheus, Victoria‑Metrics, Elasticsearch, Grafana, and Kibana.
- Hands-on experience building and implementing Prometheus pipelines, including exporters, scraping configurations, relabelling, metric routing, and integrations with long‑term storage (e.g., Victoria‑Metrics).
- Experience building and maintaining Logstash pipelines, including ingestion, parsing, filtering, enrichment, and routing of logs into Elasticsearch.
- Ability to design, build, and maintain Grafana and Kibana dashboards for metrics, logs, and performanc
📌 Lead Support Analyst (Pune)
🏢 CITIC
📍 Pune