High-Performance Computing | Cluster Administration (Bangalore Metropolitan Area)

High-Performance Computing | Cluster Administration (Bangalore Metropolitan Area)

14 Aug
|
EduBloom Talent Solutions
|
Bangalore Metropolitan Area

14 Aug

EduBloom Talent Solutions

Bangalore Metropolitan Area

HPC Administrator

High-Performance Computing | Cluster Administration

Experience

5+ years in HPC administration

Core Focus

Large-scale compute clusters for scientific, engineering and AI workloads

Environment

Linux (RHEL/SLES/Ubuntu), Slurm/PBS Pro, InfiniBand, HPE ProLiant Gen10/Gen11

About the Role

We're looking for a skilled HPC Administrator to design, deploy, maintain and optimise large-scale compute clusters powering advanced scientific, engineering and AI workloads. You'll manage high-density compute nodes and InfiniBand interconnects, run job scheduling environments, and keep the cluster ecosystem highly available, scalable and efficient.

Key Responsibilities

Linux Administration

- Comfortable working across shell environments such as BASH, sh, tcsh and ksh
- Experience administering Linux, including boot process, file systems and kernel modules
- Familiar with core OS services: SSH, telnet, FTP, NFS, DNS, DHCP, Samba and LDAP
- Working knowledge of networking tools like ping, tracert, tracemon and tcpdump
- Solid understanding of the OSI model and related networking concepts

Cluster Administration

- Install, configure and maintain Linux-based compute and login nodes
- Manage workload schedulers such as Slurm and PBS Pro, optimising job throughput and queue performance
- Oversee software provisioning, OS image deployment and firmware updates across thousands of nodes
- Manage HPC frameworks like HPE Performance Cluster Manager (HPCM), Bright Cluster Manager or equivalent

Networking & Interconnect





- Monitor and tune high-speed fabrics such as InfiniBand (HDR/NDR), including subnet managers (UFM, OpenSM) and link diagnostics
- Ensure low-latency, congestion-free data paths across compute and storage

Monitoring & Automation

- Implement cluster monitoring using tools like Grafana, Prometheus, Nagios or Opsramp
- Build automation and orchestration scripts (Bash, Python, Ansible) for provisioning, health checks and performance tuning
- Proactively identify performance bottlenecks and recommend fixes

Security & Compliance

- Maintain secure authentication, role-based access and audit controls across nodes
- Apply patches, enforce configuration management and ensure compliance with organisational IT policy

What You Bring

Essential

- Bachelor's degree in Computer Science, Engineering or a related technical field
- Strong Linux systems administration experience (RHEL/SLES/Ubuntu) in an HPC or data centre workplace
- Minimum 5 years of experience in HPC administration
- Hands-on expertise with Slurm and PBS Pro job schedulers
- Familiarity with InfiniBand networking
- Proficiency in scripting (Bash, Python or Ansible) and version control (Git)
- Good knowledge of HPE ProLiant Gen10 or Gen11 servers is preferred

Nice to Have

- Experience with containerised HPC (Singularity/Apptainer/Docker) and Kubernetes integration
- Basic understanding of VMware
- Familiarity with Lustre or GPFS filesystems
- Certifications such as RHCSA/RHCE, CKA, or vendor-specific HPC administration credentials

📌 High-Performance Computing | Cluster Administration (Bangalore Metropolitan Area)
🏢 EduBloom Talent Solutions
📍 Bangalore Metropolitan Area

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: high-performance computing | cluster administration (bangalore metropolitan area) / bangalore metropolitan area

Subscribe to this job alert:

Get the latest job offers by email for: high-performance computing | cluster administration (bangalore metropolitan area) / bangalore metropolitan area