Senior Deep Learning Algorithms Engineer BioNeMo (Mumbai)

Senior Deep Learning Algorithms Engineer BioNeMo (Mumbai)

05 Aug
|
NVIDIA
|
Mumbai

05 Aug

NVIDIA

Mumbai

Senior Deep Learning Algorithms Engineer - BioNeMo

Experience: Not Available to Not Available years

Location: Ho Chi Minh City, Vietnam

Skills: Deep Learning, Machine Learning, LLMs, VLMs, TensorRT, TensorRT-LLM, CUDA, Triton, Python, C++, PyTorch, TensorFlow, FP8, INT8, GPU performance engineering, Nsight, roofline analysis

What you will be doing:
Integrate TensorRT-LLM for BioNeMo models (Boltz1–2, OpenFold2–3) and upcoming structural biology models (RFDiffusion, DiffDock, ProteinNMN, Evo2, ESM3). Optimize models for low-latency, high-throughput inference using parallelism, quantization (FP8/INT8), and sparsity/pruning. Profile and debug deep learning workloads on GPUs, resolving kernel/graph bottlenecks in training/inference, including custom operators. Develop and validate custom GPU kernels (CUDA, Triton) for hot paths, memory-bound ops, and non-standard blocks in structural biology models. Collaborate with research to align model architecture and training with deployment constraints for smooth production transition.

What we want to see:
MS/PhD in CS, EE, Comp. Eng., or equivalent practical experience. 5+ years professional experience in deep learning/applied ML, with a track record of deploying optimized models/inference paths in production (not research prototypes). Strong foundation in transformer/diffusion architectures; direct experience with LLMs, VLMs, or large biology models (e.g., structure prediction). Proficient in PyTorch (and/or TensorFlow) for production-grade model building, debugging, and deployment.



Strong Python/C++; ability to read/modify performance-critical C++/CUDA code for inference stacks and custom ops. Practical experience with TensorRT/TensorRT-LLM: model conversion, optimization, deployment, and performance measurement (latency/throughput) under realistic conditions. Familiarity with GPU performance engineering: profiling (Nsight), roofline analysis, and optimization of kernels/memory access; experience writing/extending custom GPU kernels for model hot paths is required.

Ways to stand out from the crowd:
Led or significantly contributed to large-scale LLM/VLM/biology model serving (strict SLOs, high QPS, multi-GPU/node inference, cost/perf ownership). Deep customization of, or substantial contributions to, TensorRT-LLM, vLLM, SGLang, or comparable stacks, including debugging and extending for novel architectures. End-to-end ownership of FP8/INT8 (or other formats), including calibration, regression testing, and documenting accuracy vs. speed tradeoffs on biology workloads. Robust familiarity with protein structure, docking, or diffusion-based design and model families (e.g., OpenFold, Boltz, ESM, RFDiffusion, DiffDock)—demonstrated by benchmarks, publications, or open-source work. Repeated success taking non-text architectures (geometric, multimodal, structure-centric) from research/checkpoint to optimized, production-ready inference with clear metrics as well as examples of writing, maintaining, or upstreaming custom kernels or fused ops that produced measurable gains on real models or hardware.

📌 Senior Deep Learning Algorithms Engineer BioNeMo (Mumbai)
🏢 NVIDIA
📍 Mumbai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior deep learning algorithms engineer bionemo (mumbai) / mumbai

Subscribe to this job alert:

Get the latest job offers by email for: senior deep learning algorithms engineer bionemo (mumbai) / mumbai