- Deploy and manage ML/LLM models in production environments.
- Design and execute model load testing and performance benchmarking (latency, throughput, memory,
cost).
- Build and optimize multi-node, multi-GPU training pipelines.
- Configure and tune distributed training frameworks (data parallelism, model parallelism, pipeline parallelism).
- Optimize GPU utilization, memory footprint, and inference costs.
- Set up CI/CD pipelines for model deployment and retraining.
- Troubleshoot GPU, networking, and performance bottlenecks.
- Work across cloud platforms to ensure portability and vendor-agnostic deployments.
Required Skill Sets
- Robust experience with GPU workloads (NVIDIA GPUs, CUDA concepts).
- Proven expertise in model deployment on AWS, Azure, and GCP.
- Hands-on experience deploying models up to 20B parameters.
- Experience with distributed training (multi-node, multi-GPU setups).
- Deep understanding of load testing, stress testing, and benchmarking ML systems.