05 Sep
|
Cognizant
|
Pune
Role Overview
We are seeking seasoned professionals with deep expertise in operating and managing High-Performance Computing (HPC) platforms. The ideal candidate will have hands-on experience in designing, deploying, and maintaining HPC clusters, storage systems, and networking infrastructure, leveraging industry-leading tools and technologies.
Key Responsibilities
- HPC Infrastructure Management
- Operate and maintain HPC clusters based on CentOS, RHEL, and hardware platforms like HPE and NVIDIA DGX.
- Ensure optimal performance, scalability, and reliability of compute resources.
- Storage Administration
- Manage large-scale storage systems including Dell Isilon, VAST Storage, Lustre, and GPFS.
- Implement data lifecycle management and optimize storage performance for HPC workloads.
- Networking
- Configure and maintain InfiniBand-based networking for low-latency, high-bandwidth communication.
- Troubleshoot network performance issues and ensure secure connectivity.
- Cluster and Job Scheduling
- Administer cluster management tools such as Bright Cluster Manager, Altair Grid Manager, and IBM LSF.
- Optimize job scheduling and resource allocation for diverse workloads.
- Monitoring and Automation
- Implement monitoring solutions using Zabbix, Grafana, and ELK Stack.
- Automate provisioning and configuration using Cobbler, Chef, Ansible, and AWS ParallelCluster.
- Performance Tuning & Troubleshooting
- Conduct performance benchmarking and tuning for HPC workloads.
- Diagnose and resolve hardware/software issues across compute, storage, and network layers.
- Security & Compliance
- Ensure HPC setting adheres to security best practices and compliance standards.
Required Skills & Qualifications
- Technical Expertise
- Strong knowledge of Linux OS (CentOS, RHEL) and HPC hardware platforms (HPE, NVIDIA DGX).
- Hands-on experience with parallel file systems (Lustre, GPFS) and enterprise storage solutions.
- Proficiency in InfiniBand networking and high-speed interconnects.
- Familiarity with job schedulers and cluster management tools (IBM LSF, Bright Cluster Manager, Altair Grid Manager).
- Automation & Scripting
- Expertise in Ansible, Chef, Cobbler, and scripting languages (Bash, Python).
- Experience with AWS ParallelCluster or similar cloud-based HPC solutions.
- Monitoring & Logging
- Practical experience with Zabbix, Grafana, and ELK Stack for system health and performance monitoring.
- Soft Skills
- Strong problem-solving and analytical skills.
- Ability to work in a fast-paced environment and lead technical teams.
- Excellent communication and documentation skills.
Preferred Qualifications
- Exposure to AI/ML workloads on HPC clusters.
- Experience with containerization (Docker, Singularity) in HPC environments.
- Knowledge of security hardening for HPC systems.
Education
- Bachelors or Master’s degree in Computer Science, Engineering, or related field.
📌 Cognizant Hiring For Sr. HPC Engineer (Pune)
🏢 Cognizant
📍 Pune