Key Responsibilities
Deploy, configure and maintain HPC clusters.
Install and administer RHEL/Linux operating systems.
Configure and manage SLURM workload manager, including partitions, queues, users and job scheduling.
Work with xCAT for cluster provisioning and node management.
Install, configure and troubleshoot HPC applications.
Support CAE/CFD applications and engineering simulation workloads.
Perform cluster health checks, monitoring, troubleshooting and performance tuning.
Support HPC benchmarking and application performance analysis.
Troubleshoot compute, networking, storage and software-related issues.
Participate in customer site deployments, commissioning and technical support.
Required Skills
Linux
SLURM
xCAT
HPC cluster administration
Basic understanding of MPI / parallel computing
Knowledge of CAE & CFD applications, such as:
ANSYS
Abaqus
OpenFOAM
Altair HyperWorks
Good troubleshooting and analytical skills
Good to Have
Knowledge of InfiniBand / RoCE / high-speed networking
HPC storage technologies such as Lustre / BeeGFS
Bash/Python scripting
Experience with GPU-based HPC environments
Understanding of HPC benchmarking and performance optimization
Experience
2–5 years of relevant experience in HPC, Linux/Cluster Administration or related technologies.
? Location: Pune
? Employment Type: Full time
📌 HPC Engineer (Pune)
🏢 ByteView
📍 Pune