07 Aug
|
Profex Tech
|
Noida
Key Responsibilities:
- Deploy, manage, and optimize ML/LLM models (up to 20B parameters) in production.
- Design and implement scalable MLOps pipelines for training, deployment, and monitoring.
- Build and optimize multi-node, multi-GPU distributed training environments.
- Configure distributed training frameworks and optimize GPU utilization.
- Perform model load testing, stress testing, benchmarking, and inference optimization.
- Develop CI/CD pipelines for ML model deployment and retraining.
- Monitor, troubleshoot, and resolve infrastructure, GPU, networking, and performance bottlenecks.
- Ensure cloud-agnostic deployments across AWS, Azure, and GCP.
- Collaborate with AI/ML teams to improve model reliability, scalability, and operational efficiency.
Required Skills:
- 4+ years of experience in DevOps/MLOps/AIOps.
- Solid hands-on experience with Kubernetes, Docker, Linux, and CI/CD.
- Experience with NVIDIA GPUs, CUDA, GPU optimization, and distributed training.
- Hands-on experience with PyTorch and ML model deployment.
- Expertise in AWS, Azure, and GCP.
- Experience with Infrastructure as Code (Terraform preferred).
- Knowledge of DeepSpeed, FSDP, Megatron-LM, NCCL, or similar distributed training frameworks.
- Experience with vLLM, Triton Inference Server, TensorRT, or ONNX Runtime is preferred.
- Strong scripting skills in Python and Shell.
- Experience with monitoring tools such as Prometheus, Grafana, or OpenTelemetry.
Preferred Skills:
- Experience deploying and managing Large Language Models (LLMs).
- Knowledge of Hugging Face ecosystem and Generative AI.
- Understanding of model serving, inference optimization, autoscaling, and cost optimization.
- Exposure to ML observability and AI infrastructure.
📌 DevOps (AI+MLOps) Engineer ||H || Immediate Joiner || Gurugram (Noida)
🏢 Profex Tech
📍 Noida