10 Sep
|
Mphasis
|
Bengaluru
Skill - HPC, Slurm, ClearML, Linux
Location - Bengaluru
Experience -7-10yrs
Early joiners are preferred
Job Summary
We are seeking a highly skilled Senior Principal Infrastructure Engineer to administer and optimize our ClearML Server and associated infrastructure. The ideal candidate will have a strong background in MLOps platforms, HPC execution environments, and containerized solutions, with a focus on supporting GPU-based AI/ML workloads. This role requires a proactive approach to managing resources, ensuring security, and automating processes to enhance operational efficiency.
Responsibilities
- Administer ClearML Server, including management of agents, execution queues, projects, users, roles, experiment tracking, pipelines, datasets, artifacts, and model registry.
- Configure ClearML Agents on CPU and GPU worker nodes, integrating with HPC execution platforms such as Slurm, PBS Skilled, or Kubernetes.
- Support GPU-based AI/ML workloads utilizing NVIDIA drivers, CUDA, NCCL, UCX, and containerized environments.
- Maintain secure container execution using technologies like Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes.
- Implement confidential-computing controls leveraging AMD SEV-SNP, Intel TDX, NVIDIA Confidential Computing, Secure Boot, TPM, and remote attestation.
- Integrate authentication, RBAC, TLS certificates, secrets management, and audit controls for ClearML and confidential workloads.
- Monitor ClearML services, agents, queues, GPU utilization, task failures, scheduler integration, and overall platform health.
- Automate deployment, configuration, monitoring, and troubleshooting processes using Python, Bash, Ansible, and Git.
Mandatory Skills:
- Strong Linux administration skills, particularly with RHEL, SLES, Rocky Linux, or Ubuntu.
- Hands-on experience with ClearML administration or a comparable MLOps platform.
- Familiarity with HPC schedulers such as Slurm, PBS Professional, or LSF.
- Knowledge of GPU platforms, including NVIDIA drivers, CUDA, and distributed training basics.
- Proficiency in container technologies: Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes.
- Strong scripting and automation skills in Python, Bash, Ansible, and Git.
- Understanding of security fundamentals, including RBAC, IAM, TLS, certificates, secrets management, secure boot, and audit logging.
- Familiarity with confidential-computing concepts such as TEE, encrypted memory, TPM, remote attestation, AMD SEV-SNP, Intel TDX, or equivalent technologies.
📌 Sr. Principal Infrastr Eng (HPC) (Bengaluru)
🏢 Mphasis
📍 Bengaluru