Yobitel is building the AI-native compute layer that the next decade of AI runs on. We own and operate our own GPU clusters H100s today, B200s next. Were hiring a senior systems engineer to keep them quick, healthy, and full.
Youll own
- Cluster commissioning, hardware-level diagnostics, NUMA + topology tuning.
- Inference and training-side performance regressions: catch them at PR-time, not in production.
- The on-call rotation that pages when a rack drops or NVLink degrades.
You probably are
- Comfortable in the kernel boundary NVIDIA driver versions, CUDA toolkit pinning, RDMA fabric debugging.
- Allergic to flaky tests and "works on my box" stories.
- Happy to write a runbook the second time you do the same thing manually.