24 Sep
|
Q1 Technologies India PVT
|
Hyderabad
24 Sep
Q1 Technologies India PVT
Hyderabad
Skills: HPC AI System Management
Experience Required: 7+Years
Location: Hyderabad (Adibatla), Delhi, Pune
Topic of Evaluation / Skills*
Mandatory / Non-mandatory
Manage and maintain HPC cluster infrastructure.
Yes
Administer Linux-based compute nodes and management nodes.
Yes
Configure cluster expansion, node replacement, and hardware upgrades.
Yes
Perform OS patching and vulnerability remediation activities.
Yes
Ensure cluster availability, performance, and capacity planning.
Yes
Manage cluster health monitoring and proactive troubleshooting.
Yes
Administer NVIDIA GPU environments.
Yes
Monitor GPU utilization, temperature, memory consumption, and health.
Yes
Support AI/ML and CUDA workloads.
Yes
Configure GPU drivers and CUDA toolkits.
Yes
Troubleshoot GPU hardware and performance issues.
Yes
Support GPU lifecycle replacement and upgrades.
Yes
Manage CPU compute nodes.
Yes
Troubleshoot hardware failures.
Yes
Perform firmware and BIOS upgrades.
Yes
Capacity planning for CPU-intensive applications.
Yes
Support cluster scaling initiatives.
Yes
Experience with one or more: Slurm , PBS Pro , LSF, Torque
Yes
Manage OneFS clusters.
Yes
Storage provisioning.
Yes
Capacity management.
Yes
NFS/SMB exports.
Yes
Data protection and backup integration.
Yes
Snapshot management.
Yes
Performance tuning.
Yes
Storage health monitoring.
Yes
Hardware replacement coordination.
Yes
Job Requirements and Responsibilities
Section
Details / Example Content
Job Requirements*
Experience managing HPC clusters (CPU/GPU environments).
Knowledge of HPC schedulers such as Slurm, PBS, or LSF.
Experience with vulnerability management and patching.
Knowledge of storage and parallel file systems (GPFS/IBM Spectrum Scale preferred).
Experience in server hardware management and troubleshooting.
Understanding of performance monitoring and capacity management.
Solid troubleshooting and customer communication skills.
Key Responsibilities*
Manage and support HPC cluster infrastructure on a daily basis.
Monitor cluster health, performance, CPU/GPU utilization, and storage.
Perform vulnerability remediation and security patching.
Troubleshoot Linux, hardware, storage, and scheduler-related issues.
Coordinate hardware upgrades and replacements.
Maintain system documentation, SOPs, and procedures.
Support users and application teams using the HPC environment.
Participate in incident, problem, and change management activities.
Ensure SLA compliance and timely resolution of issues.
Work with vendors and internal teams for complex issues.
Job Qualifications & Skills
Section
Details / Example Content
Domain
Manufacturing
Soft Skills
- Excellent communication
- Team collaboration
- Documentation and knowledge sharing
Education Requirements
Bachelor's/master's or equivalent (can be marked optional or flexible)
Certifications
Linux Certifications
Red Hat Certified System Administrator (RHCSA)
Red Hat Certified Engineer (RHCE)
Linux Professional Institute Certification (LPIC-1 / LPIC-2)
HPC-Specific Certifications
NVIDIA Certified Professional (GPU & AI Infrastructure)
NVIDIA DGX / GPU Platform Training Certifications
Intel HPC and Parallel Computing Courses
📌 HPC Administrator (Hyderabad)
🏢 Q1 Technologies India PVT
📍 Hyderabad