08 Aug
|
Profex Tech
|
India
Key Roles & Responsibilities
• Deploy and manage ML/LLM models in production settings.
• Design and execute model load testing and performance benchmarking (latency, throughput, memory,
cost).
• Build and optimize multi-node, multi-GPU training pipelines.
• Configure and tune distributed training frameworks (data parallelism, model parallelism, pipeline
parallelism).
• Optimize GPU utilization, memory footprint, and inference costs.
• Set up CI/CD pipelines for model deployment and retraining.
• Troubleshoot GPU, networking, and performance bottlenecks.
• Work across cloud platforms to ensure portability and vendor-agnostic deployments.
Required Skill Sets
• Solid experience with GPU workloads (NVIDIA GPUs, CUDA concepts).
• Proven expertise in model deployment on AWS, Azure, and GCP.
• Hands-on experience deploying models up to 20B parameters.
• Experience with distributed training (multi-node, multi-GPU setups).
• Deep understanding of load testing, stress testing, and benchmarking ML systems.
📌 Devops Mlops+aiops Engineer Gurugram (India)
🏢 Profex Tech
📍 India