About Gruve
Gruve is an innovative software services startup dedicated to transforming enterprises to AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions. As a well-funded early-stage startup, Gruve offers a dynamic environment with strong customer and partner networks.
Position summary:
Senior analyst and shift anchor for the SOC pod, and L2 for the Kubernetes/OpenShift-based PulseAI platform layer. Owns triage quality across the rotation, handles high-severity security incidents — including Kubernetes and container detections — through to handoff, remediates PulseAI and OpenShift platform incidents to the Standard/Premium restoration objectives (Severity 1 in 4/2 hours), executes platform changes and patching within maintenance windows, and manages escalations to L3 platform engineering. Infrastructure-layer faults (GPU-server hardware, nodes, fabric, switches, storage) are handed to the NOC pod and are not part of this role.
Key Roles & Responsibilities:
- Act as shift senior: final quality gate on investigations and escalations; handle P1/P2 security alerts end-to-end to L3 handoff including scoping and evidence preservation.
- Own L2 handling of Kubernetes/OpenShift security detections — kube-audit anomalies (privilege escalation, secrets access, RBAC and service-account misuse), suspicious workload and pod behaviour, image and admission-policy violations, Cilium/Hubble flow alerts — executing containment within the authority matrix (namespace isolation via NetworkPolicy, workload scale-down or cordon, credential/token revocation) and feeding tuning back into the detection backlog.
- Drive shift-level metrics: SLA adherence, false-positive rate, reopen rate.
- Remediate PulseAI platform incidents within the authority matrix — control-plane and platform-service recovery, authentication/SSO and RBAC faults, tenancy and quota-enforcement failures, endpoint deployment and model-serving failures, observability outages — and restore PulseAI configuration state (organisation/department/project hierarchy,
quota allocations, users and roles) from Gruve backups when required.
- Remediate OpenShift platform-side cluster incidents — operator degradation, scheduling and capacity, storage/PVC faults, image registry and ingress, RBAC and identity-provider integration, node NotReady triage to a platform-versus-hardware determination — using oc/kubectl and cluster diagnostics; hand confirmed node-hardware, cluster-network and fabric faults to the NOC pod through the ITSM handoff, and structured RCAs to L3.
- Execute PulseAI patch releases and OpenShift z-stream patches in the monthly maintenance window, and emergency security remediation within the tier window (72 hours Premium / 5 business days Standard); enforce the pre-change gate that no cluster upgrade proceeds without the customer's written confirmation of a verified backup.
- Manage platform escalations: drive L3 platform-engineering escalations and Red Hat support cases opened for platform defects to closure, keep the customer informed at the SLA cadence,and pause/resume restoration clocks correctly when waiting on the customer, a vendor or a change approval.
- Own platform observability hygiene: Grafana dashboards, alert thresholds recorded in the SLA appendix (performance degradation, capacity), and runbooks for recurring platform faults and Kubernetes security use cases.
- Coach and quality-review the L1 security analysts; own runbook accuracy for the SOC pod.
Mandatory Qualifications:
- BE/BTech (CS/IT/E&TC;) or equivalent.
- 4–6 years SOC or 24×7 platform-operations experience with demonstrable incident handling and remediation ownership.
- Strong multi-source triage across cloud audit, network-security (firewall, WAF, flow) and identity telemetry.
- Hands-on triage of Kubernetes/container security alerts — kube-audit events,
workload anomalies and pod-level network flows (Cilium/Hubble) — on GKE and/or OpenShift, with a working grasp of container attack paths (MITRE ATT&CK; for Containers).
- Hands-on Red Hat OpenShift / Kubernetes administration in production — troubleshooting pods, nodes, operators, storage, RBAC and ingress with oc/kubectl; applying z-stream patches and operator updates within change control; reading control-plane and workload logs in a metrics/logging observability stack (Grafana).
- Working understanding of GPU workloads at the platform layer — NVIDIA GPU Operator as an OpenShift component, DCGM-class telemetry as read in Grafana, common GPU health signals — sufficient to separate a platform fault from a hardware fault and hand the latter to the NOC pod; transparent grasp of the monitor / remediate / escalate boundaries between platform, hardware and customer workload.
- Audit-grade documentation; ability to run a shift independently and calmly under Severity 1 pressure and to communicate status to customer authorised contacts.
- Mentoring aptitude — this role carries the shift's junior bench.
Preferred Qualifications:
- Incident-handling / forensics certification-level knowledge (e.g., GCIH, CHFI or equivalent); scripting for enrichment/automation.
- CKA/CKS or KCSA; exposure to admission controls and image/runtime security tooling.
- Red Hat OpenShift Administration certification (EX280) or RHCSA; exposure to NVIDIA AI Enterprise (NIM microservices), model-serving endpoints and GPU workload scheduling on OpenShift.
- Exposure to HashiCorp Vault, SAML 2.0 SSO / RBAC troubleshooting, and n8n or similar platform applications.
- Prior GPU-cloud, neocloud, hyperscale or data-center customer exposure.
Why Gruve
At Gruve, we foster a culture of innovation, collaboration, and continuous learning. We are committed to building a diverse and inclusive workplace where everyone can thrive and contribute their best work. If you’re passionate about technology and eager to make an impact, we’d love to hear from you.
Gruve is an equal opportunity employer. We welcome applicants from all backgrounds and thank all who apply; however, only those selected for an interview will be contacted.
📌 Security Analyst II (Pune)
🏢 Gruve
📍 Pune