About Gruve
Gruve is an innovative software services startup dedicated to transforming enterprises to AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions. As a well-funded early-stage startup, Gruve offers a dynamic workplace with strong customer and partner networks.
Position Summary
Entry-level engineer on the 24×7 NOC monitoring rotation. First responder to alerts from the AI Fabrik network estate — data-center fabric, edge routers/firewalls and out-of-band console management — and to PulseAI infrastructure health alerts (GPU servers, control-plane/infrastructure nodes, OpenShift cluster nodes, switches and storage) arriving through the Gruve outbound collector, working strictly from runbooks under the shift senior and building toward independent first-level diagnostics within 6 months.
Key Roles & Responsibilities
- Monitor the network monitoring dashboards and the PulseAI infrastructure dashboards (Grafana); acknowledge fabric, edge, firewall and infrastructure alerts within the tier SLA and perform first-level triage per documented runbooks.
- Classify alerts (actionable / informational / false alarm) with evidence and record complete investigation notes in the ITSM ticketing platform (the SLA system of record).
- Watch GPU-server, control-plane/infrastructure node, switch and storage health signals — device reachability, interface and link errors, optics, CPU/memory, GPU utilisation/temperature, OpenShift node status, storage capacity thresholds, telemetry reachability — log deviations and escalate per the severity matrix with accurate context and timelines.
- Recognise PulseAI infrastructure severity conditions — node or GPU loss, front-end network unreachable, back-end RoCEv2 fabric degraded, out-of-band management unreachable,
📌 Network Analyst I (Pune)
🏢 Gruve
📍 Pune