About Gruve
Gruve is an innovative software services startup dedicated to transforming enterprises to AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions. As a well-funded early-stage startup, Gruve offers a energetic environment with strong customer and partner networks.
Position summary:
Shift engineer owning monitoring and first-fix for the AI Fabrik network estate — EVPN-VXLAN data-center fabric, edge routers/firewalls, next-generation firewall HA pairs and out-of-band console management — and L1 for the PulseAI infrastructure layers: monitoring GPU servers, control-plane/infrastructure nodes, the front-end network and optional RoCEv2 back-end fabric, supported customer switches and storage, with first-level diagnostics and vendor case creation.
Key responsibilities:
- Monitor and first-fix fabric,
edge and firewall alerts; run structured troubleshooting before escalation.
- Execute standard changes under change control: port turn-ups, ACL updates, code upgrades in maintenance windows.
- Own device configuration hygiene: backup verification, drift checks, golden-config compliance reporting.
- Monitor PulseAI customer infrastructure via the collector: GPU-server and node health (availability, GPU utilisation/thermal, NIC and link errors, out-of-band management reachability), OpenShift node and cluster-network status, front-end network reachability of every cluster node, back-end RoCEv2 fabric health, switch telemetry (SNMP/syslog/streaming) and storage capacity/health; acknowledge within the tier SLA.
- Perform first-level diagnostics to separate network-side, hardware-side, storage-side and platform-side faults; confirm hardware faults and open the vendor case within the tier window (60/30 minutes), record the case reference and track to
📌 Network Consultant I (Pune)
🏢 Gruve
📍 Pune