29 Aug
|
Cognizant
|
Hyderabad
29 Aug
Cognizant
Hyderabad
Role Overview
Weare seeking seasoned professionals with deep expertise in operating andmanaging High-Performance Computing (HPC) platforms. The ideal candidatewill have hands-on experience in designing, deploying, and maintaining HPCclusters, storage systems, and networking infrastructure, leveragingindustry-leading tools and technologies.
KeyResponsibilities
HPCInfrastructure Management o Operate and maintain HPCclusters based on CentOS, RHEL, and hardware platforms like HPEand NVIDIA DGX.
o Ensure optimal performance,scalability, and reliability of compute resources.
StorageAdministration o Manage large-scale storagesystems including Dell Isilon, VAST Storage, Lustre, and GPFS.
o Implement data lifecyclemanagement and optimize storage performance for HPC workloads.
Networking o Configure and maintain InfiniBand-basednetworking for low-latency, high-bandwidth communication.
o Troubleshoot networkperformance issues and ensure secure connectivity.
Cluster andJob Scheduling o Administer clustermanagement tools such as Bright Cluster Manager, Altair Grid Manager,and IBM LSF.
o Optimize job scheduling andresource allocation for diverse workloads.
Monitoringand Automation o Implement monitoringsolutions using Zabbix, Grafana, and ELK Stack.
o Automate provisioning andconfiguration using Cobbler, Chef, Ansible, and AWSParallelCluster.
PerformanceTuning & Troubleshooting o Conduct performancebenchmarking and tuning for HPC workloads.
o Diagnose and resolvehardware/software issues across compute, storage, and network layers.
Security& Compliance o Ensure HPC environmentadheres to security best practices and compliance standards.
RequiredSkills & Qualifications
TechnicalExpertise o Strong knowledge of LinuxOS (CentOS, RHEL) and HPC hardware platforms (HPE, NVIDIA DGX).
o Hands-on experience with parallelfile systems (Lustre, GPFS) and enterprise storage solutions.
o Proficiency in InfiniBandnetworking and high-speed interconnects.
o Familiarity with jobschedulers and cluster management tools (IBM LSF, Bright Cluster Manager,Altair Grid Manager).
Automation& Scripting o Expertise in Ansible,Chef, Cobbler, and scripting languages (Bash, Python).
o Experience with AWSParallelCluster or similar cloud-based HPC solutions.
Monitoring& Logging o Practical experience with Zabbix,Grafana, and ELK Stack for system health and performancemonitoring.
Soft Skills o Solid problem-solving andanalytical skills.
o Ability to work in afast-paced environment and lead technical teams.
o Excellent communication anddocumentation skills.
PreferredQualifications
Exposure to AI/MLworkloads on HPC clusters.
Experiencewith containerization (Docker, Singularity) in HPC environments.
Knowledge of securityhardening for HPC systems.
Education
Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
📌 Sr. HPC ENGINEER (Hyderabad)
🏢 Cognizant
📍 Hyderabad