07 Aug
|
Talworx Solutions
|
Bengaluru
07 Aug
Talworx Solutions
Bengaluru
About the role:
We are seeking a highly skilled High-Performance Computing (HPC)
Administrator to design, deploy, maintain, and optimize large-scale compute
clusters which could be used for advanced scientific, engineering and AI workloads.
You will be responsible for managing high-density compute nodes, InfiniBand
interconnects, job scheduling environments, ensuring high availability, scalability,
and efficiency across the cluster ecosystem.
Key Responsibilities
Linux Skills
• Familiarity with one, or more, shell environments (BASH, sh, tcsh, ksh, etc.)
• Experience administering Linux (boot process, file systems, kernel modules)
• Familiarity with core OS services such as: SSH, telnet, FTP, NFS, DNS,
DHCP, Samba, LDAP
• Familiarity with networking tools (ping, tracert, tracemon, tcpdump, etc.)
• Understanding of the OSI model and related concepts
Cluster Administration
• Install, configure, and maintain Linux-based compute and login nodes.
• Manage workload schedulers (e.g., Slurm, PBS Pro) and optimize job
throughput and queue performance.
• Oversee software provisioning, OS image deployment, and firmware updates
across thousands of nodes.
• Manage HPC management frameworks such as HPE Performance Cluster
Manager (HPCM), Bright Cluster Manager, or equivalent.
Networking & Interconnect
• Monitor and tune high-speed fabrics such as InfiniBand (HDR/NDR)
including subnet managers (UFM, OpenSM) and link diagnostics.
• Need to ensure low latency,
congestion-free data paths across compute and
storage.
Monitoring & Automation
• Implement comprehensive cluster monitoring using tools like Grafana,
Prometheus, Nagios or Opsramp.
• Develop automation and orchestration scripts (Bash, Python, Ansible) for
provisioning, health checks, and performance tuning.
• Proactively identify performance bottlenecks and propose corrective actions.
Security & Compliance
• Maintain secure authentication, role-based access, and audit controls across
nodes.
• Apply patches, enforce configuration management, and ensure compliance
with organizational IT policies.
Qualifications
Essential:
• Bachelors degree in Computer Science, Engineering or a related technical
field.
• Strong experience in Linux systems administration (RHEL/SLES/Ubuntu)
in an HPC or data center workplace.
• Minimum 5 years of experience in HPC administration is valuable.
• Hands-on expertise with Slurm, PBS Pro job schedulers.
• Familiarity with InfiniBand networking
• Proficiency in scripting (Bash, Python, or Ansible) and version control (Git).
• Good knowledge about the HPE Proliant Gen 10 or Gen11 servers is
preferable.
Desirable:
• Experience with containerized HPC (Singularity/Apptainer/Docker) and
Kubernetes integration.
• Basic understanding of VMware is good.
• Familiarity with Lustre/GPFS filesystems,
• Certifications like RHCSA/RHCE, or CKA or vendor-specific HPC
administration credentials.
📌 High-Performance Computing (HPC) Administrator (Bengaluru)
🏢 Talworx Solutions
📍 Bengaluru