30 Jul
|
Profex Tech
|
India
Key Responsibilities:
Deploy, manage, and optimize ML/LLM models (up to 20B parameters) in production.
Design and implement scalable MLOps pipelines for training, deployment, and monitoring.
Build and optimize multi-node, multi-GPU distributed training environments.
Configure distributed training frameworks and optimize GPU utilization.
Perform model load testing, stress testing, benchmarking, and inference optimization.
Develop CI/CD pipelines for ML model deployment and retraining.
Monitor, troubleshoot, and resolve infrastructure, GPU, networking, and performance bottlenecks.
Ensure cloud-agnostic deployments across AWS, Azure, and GCP.
Collaborate with AI/ML teams to improve model reliability, scalability, and operational efficiency.
Required Skills:
4+ years of experience in DevOps/MLOps/AIOps.
Robust hands-on experience with Kubernetes, Docker, Linux, and CI/CD.
Experience with NVIDIA GPUs, CUDA, GPU optimization, and distributed training.
Hands-on experience with PyTorch and ML model deployment.
Expertise in AWS, Azure, and GCP.
Experience with Infrastructure as Code (Terraform preferred).
Knowledge of DeepSpeed, FSDP, Megatron-LM, NCCL, or similar distributed training frameworks.
Experience with vLLM, Triton Inference Server, TensorRT, or ONNX Runtime is preferred.
Robust scripting skills in Python and Shell.
Experience with monitoring tools such as Prometheus, Grafana, or OpenTelemetry.
Preferred Skills:
Experience deploying and managing Large Language Models (LLMs).
Knowledge of Hugging Face ecosystem and Generative AI.
Understanding of model serving, inference optimization, autoscaling, and cost optimization.
Exposure to ML observability and AI infrastructure.
📌 Devops Ai+mlops Engineer H Immediate Joiner Gurugram Delhi (India)
🏢 Profex Tech
📍 India