24 Sep
|
HL Mando Softtech
|
Gurugram
24 Sep
HL Mando Softtech
Gurugram
MLOps Engineer
Experience : 3+ years
Role Overview
We are looking for an MLOps Engineer with strong hands-on experience in DevOps, Kubernetes, CI/CD, infrastructure automation, and SRE practices.
The role will focus primarily on building and operating reliable ML infrastructure, deploying applications and ML workloads, automating infrastructure and deployments, and maintaining production systems and troubleshooting.
Key Responsibilities
- Design, deploy, maintain, and troubleshoot Kubernetes clusters and workloads.
- Manage containerized applications using Docker, Kubernetes, Helm, and containerd.
- Build and maintain CI/CD and GitOps pipelines for applications and ML workloads.
- Automate infrastructure and operational tasks using Python/Bash and tools such as Ansible.
- Deploy and manage ML training and inference workloads on Kubernetes.
- Manage GPU-based workloads and troubleshoot GPU/container runtime issues.
- Implement production monitoring, logging, alerting, and observability.
- Perform incident troubleshooting, root-cause analysis, capacity planning, and reliability improvements.
- Implement health checks, resource management, autoscaling, rollback, backup, and recovery mechanisms.
- Troubleshoot Linux, networking, storage, containers, Kubernetes, and application-level issues.
- Maintain secure and reliable production infrastructure.
- Work closely with ML Engineers and Software Engineers to productionize ML models.
Required Skills
- 3+ years in MLOps, DevOps, SRE, Platform Engineering, Infrastructure Engineering.
- Robust hands-on Kubernetes experience.
- Strong Linux administration and troubleshooting.
- Docker / containerization.
- CI/CD and GitOps concepts.
- Python and Bash scripting.
- Networking fundamentals: DNS, TCP/IP, HTTP/HTTPS, ingress, load balancing.
- Monitoring and observability.
- Production troubleshooting and incident management.
- Understanding of infrastructure reliability, availability, and performance.
- Infrastructure solutions planning
Good to Have
MLOps
- Kubeflow, MLflow
- Distributed Training, Distributed Inferencing, Model Management, Disaster Recovery
Data
- DataOps, Disaster Recovery, Scaling
- MinIO / S3
- Kafka, Apache Airflow
- PostgreSQL / MySQL
Cloud
- AWS / Azure / GCP
- Hybrid Architecture.
Infrastructure
- Terraform
- Ansible
- Harbor
- Prometheus / Grafana
- Loki / ELK
- Istio / cert-manager
Ideal Candidate
- Strong DevOps/SRE mindset with excellent troubleshooting skills.
- Comfortable operating production Kubernetes environments.
- Strong understanding of Linux, networking, containers, and infrastructure.
- Automation-oriented and able to reduce manual operational work.
- Comfortable working with on-premise/local infrastructure and GPU servers.
- Able to learn and work across ML platform technologies.
📌 Hiring For ML OPs Engineer (Gurugram)
🏢 HL Mando Softtech
📍 Gurugram