Design and operate large-scale GPU clusters running NVIDIA H100 and B200 hardware for AI inference and training workloads.
Engineering Remote / Hyderabad full time Remote-eligible
About the role
Yobitel is building the AI-native compute layer that the next decade of AI runs on. We own and operate our own GPU clusters — H100s today, B200s next. We're hiring a senior systems engineer to keep them rapid, healthy, and full.
You'll own
- Cluster commissioning, hardware-level diagnostics, NUMA + topology tuning.
- Inference and training-side performance regressions: catch them at PR-time, not in production.
- The on-call rotation that pages when a rack drops or NVLink degrades.
You probably are
- Comfortable in the kernel boundary — NVIDIA driver versions, CUDA toolkit pinning, RDMA fabric debugging.
- Allergic to flaky tests and "works on my box" stories.
- Happy to write a runbook the second time you do the same thing manually.
Role reference: YJ-00001
📌 Senior GPU Systems Engineer (India)
🏢 Yobitel Communications
📍 India
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.