27 Aug
|
Talworx Solutions
|
Bengaluru
27 Aug
Talworx Solutions
Bengaluru
About the role
We are seeking a highly skilled High-Performance Computing (HPC)
Administrator to design, deploy, maintain, and optimize large-scale compute clusters which could be used for advanced scientific, engineering and AI workloads.
You will be responsible for managing high-density compute nodes, InfiniBand interconnects, job scheduling environments, ensuring high availability, scalability,
and efficiency across the cluster ecosystem.
Key Responsibilities
Linux Skills
- Familiarity with one, or more, shell environments (BASH, sh, tcsh, ksh, etc.)
- Experience administering Linux (boot process, file systems, kernel modules)
- Familiarity with core OS services such as: SSH, telnet, FTP, NFS, DNS,
DHCP, Samba, LDAP
- Familiarity with networking tools (ping, tracert, tracemon, tcpdump, etc.)
- Understanding of the OSI model and related concepts
Cluster Administration
- Install, configure, and maintain Linux-based compute and login nodes.
- Manage workload schedulers (e.g., Slurm, PBS Pro) and optimize job throughput and queue performance.
- Oversee software provisioning, OS image deployment, and firmware updates across thousands of nodes.
- Manage HPC management frameworks such as HPE Performance Cluster
Manager (HPCM), Bright Cluster Manager, or equivalent.
Networking & Interconnect
- Monitor and tune high-speed fabrics such as InfiniBand (HDR/NDR)
including subnet managers (UFM, OpenSM) and link diagnostics.
- Need to ensure low latency,
congestion-free data paths across compute and storage.
Monitoring & Automation
- Implement comprehensive cluster monitoring using tools like Grafana,
Prometheus, Nagios or Opsramp.
- Develop automation and orchestration scripts (Bash, Python, Ansible) for provisioning, health checks, and performance tuning.
- Proactively identify performance bottlenecks and propose corrective actions.
Security & Compliance
- Maintain secure authentication, role-based access, and audit controls across nodes.
- Apply patches, enforce configuration management, and ensure compliance with organizational IT policies.
Qualifications
Essential
- Bachelors degree in Computer Science, Engineering or a related technical field.
- Strong experience in Linux systems administration (RHEL/SLES/Ubuntu)
in an HPC or data center workplace.
- Minimum 5 years of experience in HPC administration is valuable.
- Hands-on expertise with Slurm, PBS Pro job schedulers.
- Familiarity with InfiniBand networking
- Proficiency in scripting (Bash, Python, or Ansible) and version control (Git).
- Good knowledge about the HPE Proliant Gen 10 or Gen11 servers is preferable.
Desirable
- Experience with containerized HPC (Singularity/Apptainer/Docker) and
Kubernetes integration.
- Basic understanding of VMware is good.
- Familiarity with Lustre/GPFS filesystems,
- Certifications like RHCSA/RHCE, or CKA or vendor-specific HPC administration credentials.
📌 High-Performance Computing (HPC) Administrator (Bengaluru)
🏢 Talworx Solutions
📍 Bengaluru