We are looking for a Senior FDE with solid expertise in AI infrastructure, GPU clusters, Kubernetes, and distributed systems to work directly with customers on production AI workloads.
Key Responsibilities
Design and optimize GPU infrastructure and LLM inference platforms.
Work with NVIDIA/AMD GPUs, CUDA/ROCm, vLLM, TensorRT-LLM, and SGLang.
Deploy and manage Kubernetes-based AI infrastructure.
Automate infrastructure using Terraform, Ansible, and Helm.
Troubleshoot GPU drivers, NCCL/RCCL, networking, and storage issues.
Collaborate directly with customer engineering teams to deliver production AI solutions.
Required Skills
6+ years of experience in AI Infrastructure, FDE, Distributed Systems, or Technical Consulting.
Robust Linux, Kubernetes, Python, and Go skills.
Hands-on experience with NVIDIA/AMD GPU infrastructure and CUDA/ROCm.
Knowledge of LLM inference, GPU optimization, RDMA/InfiniBand/RoCE.
Experience with Terraform/Helm and production cloud infrastructure.
Solid customer-facing and problem-solving skills.
Preferred: Experience with production AI systems, GPU vendors/cloud platforms, and high-performance computing environments.