14 Aug
|
EduBloom Talent Solutions
|
Bangalore Metropolitan Area
14 Aug
EduBloom Talent Solutions
Bangalore Metropolitan Area
HPC Administrator
High-Performance Computing | Cluster Administration
Experience
5+ years in HPC administration
Core Focus
Large-scale compute clusters for scientific, engineering and AI workloads
Environment
Linux (RHEL/SLES/Ubuntu), Slurm/PBS Pro, InfiniBand, HPE ProLiant Gen10/Gen11
About the Role
We're looking for a skilled HPC Administrator to design, deploy, maintain and optimise large-scale compute clusters powering advanced scientific, engineering and AI workloads. You'll manage high-density compute nodes and InfiniBand interconnects, run job scheduling environments, and keep the cluster ecosystem highly available, scalable and efficient.
Key Responsibilities
Linux Administration
- Comfortable working across shell environments such as BASH, sh, tcsh and ksh
- Experience administering Linux, including boot process, file systems and kernel modules
- Familiar with core OS services: SSH, telnet, FTP, NFS, DNS, DHCP, Samba and LDAP
- Working knowledge of networking tools like ping, tracert, tracemon and tcpdump
- Solid understanding of the OSI model and related networking concepts
Cluster Administration
- Install, configure and maintain Linux-based compute and login nodes
- Manage workload schedulers such as Slurm and PBS Pro, optimising job throughput and queue performance
- Oversee software provisioning, OS image deployment and firmware updates across thousands of nodes
- Manage HPC frameworks like HPE Performance Cluster Manager (HPCM), Bright Cluster Manager or equivalent
Networking & Interconnect
- Monitor and tune high-speed fabrics such as InfiniBand (HDR/NDR), including subnet managers (UFM, OpenSM) and link diagnostics
- Ensure low-latency, congestion-free data paths across compute and storage
Monitoring & Automation
- Implement cluster monitoring using tools like Grafana, Prometheus, Nagios or Opsramp
- Build automation and orchestration scripts (Bash, Python, Ansible) for provisioning, health checks and performance tuning
- Proactively identify performance bottlenecks and recommend fixes
Security & Compliance
- Maintain secure authentication, role-based access and audit controls across nodes
- Apply patches, enforce configuration management and ensure compliance with organisational IT policy
What You Bring
Essential
- Bachelor's degree in Computer Science, Engineering or a related technical field
- Strong Linux systems administration experience (RHEL/SLES/Ubuntu) in an HPC or data centre workplace
- Minimum 5 years of experience in HPC administration
- Hands-on expertise with Slurm and PBS Pro job schedulers
- Familiarity with InfiniBand networking
- Proficiency in scripting (Bash, Python or Ansible) and version control (Git)
- Good knowledge of HPE ProLiant Gen10 or Gen11 servers is preferred
Nice to Have
- Experience with containerised HPC (Singularity/Apptainer/Docker) and Kubernetes integration
- Basic understanding of VMware
- Familiarity with Lustre or GPFS filesystems
- Certifications such as RHCSA/RHCE, CKA, or vendor-specific HPC administration credentials
📌 High-Performance Computing | Cluster Administration (Bangalore Metropolitan Area)
🏢 EduBloom Talent Solutions
📍 Bangalore Metropolitan Area