18 Aug
|
Bellcom Technologies
|
Bhopal
18 Aug
Bellcom Technologies
Bhopal
Job Description
Senior MLOps & DevOps Engineer
LocationBhopal, MP (In office)Experience610 YearsReports ToLead AI EngineerOpenings1Role Overview
Own and operate the end-to-end MLOps platform for enterprise AI/ML solutions on secure on-premise and air-gapped infrastructure. This is not a traditional DevOps role hands-on production experience with GPU clusters (H100/A100), Kubernetes, LLM inference platforms (vLLM/Triton), and air-gapped deployments is mandatory. The ideal candidate can stand up the full MLOps stack from bare metal without internet access and has operated AI/ML systems in defence, government, or classified environments.
Key Responsibilities
- Design and operate the full ML pipeline: ingestion training validation deployment monitoring drift detection automated retraining.
- Build and maintain ML pipelines with MLflow, Kubeflow, and/or Airflow; manage experiment tracking, model registry, and versioning.
- Deploy and manage LLM/ML inference with vLLM, Triton, or TGI on local GPU hardware.
- Administer Kubernetes on bare-metal on-premise infrastructure; manage NVIDIA GPU Operator, Helm, RBAC, and storage integration.
- Configure and maintain NVIDIA GPU infrastructure: CUDA, cuDNN, TensorRT, NCCL, GPU drivers, and multi-GPU scheduling.
- Build CI/CD pipelines for model code; automate testing, validation, and deployment via IaC (Terraform, Ansible, Helm).
- Set up model monitoring dashboards: accuracy tracking, data drift detection, and performance alerting (Prometheus, Grafana).
- Manage air-gapped environments: offline package repos,
private container registries, RBAC, audit logging, and security controls.
- Implement HA, disaster recovery, backup, and incident management for the ML platform.
Required Skills & Experience
- 610 years in DevOps/MLOps; 3+ yrs DevOps, 2+ yrs dedicated MLOps with production GPU infrastructure.
- Hands-on NVIDIA GPU ops: H100, H200, A100, L40S or RTX Ada; CUDA, cuDNN, TensorRT, NCCL, multi-GPU clusters.
- Production experience: MLflow, Kubeflow/Airflow, Docker, Kubernetes (bare metal), CI/CD, and IaC (Terraform/Ansible/Helm).
- LLM inference platforms: vLLM, NVIDIA Triton Inference Server, Ollama, or TGI on local GPU hardware.
- Proven experience operating AI/ML in air-gapped, classified, defence, or government environments with offline registries and repos.
- Ability to provision the complete local ML stack from bare metal (GPU K8s registry inference monitoring) without internet access.
- Python automation, Linux administration, FastAPI/Flask model serving, Prometheus/Grafana monitoring.
- Security controls: RBAC, secrets management, network policies, audit logging on on-premise infrastructure.
Preferred / Good to Have
- DGX-class systems (DGX A100/H100); MinIO/Ceph distributed object storage; DVC for data versioning.
- HA/DR configuration for ML serving; RPA platform integration; model governance frameworks.
Qualifications
- B.Tech / M.Tech in Computer Science, Software Engineering, or related field.
- Kubernetes (CKA) or MLOps certifications are a plus.
📌 Senior MLOps & DevOps Engineer (Bhopal)
🏢 Bellcom Technologies
📍 Bhopal