Dive in and do the best work of your career at DigitalOcean. Journey alongside a strong community of top talent who are relentless in their drive to build the simplest scalable cloud. If you have a growth mindset, naturally like to think big and bold, and are energized by the fast-paced environment of a true industry disruptor, you’ll find your place here. We value winning together—while learning, having fun, and making a profound difference for the dreamers and builders in the world.
Position Overview
We are looking for a Senior Forward Deployed Engineer I (FDE) who is passionate about operationalizing and collaborating closely with strategic AI enterprises and high-growth startups. As AI-native startups scale, serving large generative AI models at low latency without blowing up infrastructure budgets is their biggest hurdle.
That’s where you come in. As a Sr. Forward Deployed AI Inference Engineer (FDE) I, you won't just optimize benchmarks in a lab—you will embed directly with high-growth AI founders to solve high-throughput,
cluster-scale LLM serving challenges on DigitalOcean’s GPU cloud. You operate at the intersection of distributed systems, specialized AI hardware, and real-world customer impact.
You will own the full lifecycle of high-performance LLM deployment, turning raw GPU compute into hyper-productive, resilient serving infrastructure for our high-value customers. You will spend your time architecting distributed inference systems, profiling network bottlenecks, debugging KV-cache locality issues, and shipping code that directly reduces time-to-first-token (TTFT) and time-per-output-token (TPOT) for top AI teams.
What You’ll Do
Technical Leadership: Act as an AI Inference lead on the FDE team, driving the end-to-end design, development, and delivery of critical AI workloads leveraging large generative AI models.
Design & Scale Distributed AI Inference Systems: Architect and deploy production-grade, multi-tenant LLM inference engines using Kubernetes-na