Key Responsibilities
- Deploy, configure and maintain HPC clusters.
- Install and administer RHEL/Linux operating systems.
- Configure and manage SLURM workload manager, including partitions, queues, users and job scheduling.
- Work with xCAT for cluster provisioning and node management.
- Install, configure and troubleshoot HPC applications.
- Support CAE/CFD applications and engineering simulation workloads.
- Perform cluster health checks, monitoring, troubleshooting and performance tuning.
- Support HPC benchmarking and application performance analysis.
- Troubleshoot compute, networking, storage and software-related issues.
- Participate in customer site deployments, commissioning and technical support.
Required Skills
- Linux
- SLURM
- xCAT
- HPC cluster administration
- Basic understanding of MPI / parallel computing
- Knowledge of CAE & CFD applications, such as:
- ANSYS
- Abaqus
- OpenFOAM
- Altair HyperWorks
- Good troubleshooting and analytical skills
Valuable to Have
- Knowledge of InfiniBand / RoCE / high-speed networking
- HPC storage technologies such as Lustre / BeeGFS
- Bash/Python scripting
- Experience with GPU-based HPC environments
- Understanding of HPC benchmarking and performance optimization
Experience2–5 years of relevant experience in HPC, Linux/Cluster Administration or related technologies. ? Location: Pune
? Employment Type: Full-time
📌 HPC Engineer (Pune)
🏢 ByteView
📍 Pune