02 Oct
|
Netweb Technologies India
|
Faridabad
02 Oct
Netweb Technologies India
Faridabad
Job Summary
We are looking for a Senior Engineer HPC with strong expertise in Linux server administration, enterprise server hardware, system troubleshooting, and AI/HPC infrastructure. The role involves end-to-end troubleshooting and maintenance of enterprise servers and workstations, covering hardware, firmware, RAID/storage, operating systems, networking, and application workloads.
The candidate will be responsible for diagnosing and resolving complex infrastructure issues, performing hardware replacements and firmware upgrades, managing Linux environments, monitoring system performance, and conducting Root Cause Analysis (RCA) for critical or recurring incidents. The role also requires exposure to GPU-based servers, NVIDIA GPUs/CUDA, HPC platforms such as Slurm/OpenPBS, virtualization, networking, and high-performance computing environments.
The position requires close coordination with OEMs, vendors, application/software teams, network teams, and internal technical teams to ensure timely incident resolution, SLA adherence, proper documentation, and reliable infrastructure operations.
Qualification:
B.E./B.Tech in Computer Science / Information Technology or equivalent
Key Responsibilities:
1. Server & Hardware Infrastructure
- Installation, configuration and maintenance of enterprise servers and workstations.
- Troubleshoot server hardware issues involving CPU, RAM, motherboard, RAID, HDD/SSD/NVMe, PSU, NIC, HBA and GPU.
- Perform hardware replacement, component-level troubleshooting and preventive maintenance.
- Troubleshoot POST, boot, disk, RAID, firmware, thermal and power-related issues.
- Work with server management interfaces such as IPMI/iLO/iDRAC/BMC.
- Perform BIOS, BMC, RAID controller, NIC and other firmware upgrades.
- Analyze hardware logs, SEL logs and system diagnostics to identify root cause.
2. Linux / Operating System Administration
- Installation, configuration, patching and upgrading of RHEL, Rocky/Alma Linux, Ubuntu and other Linux distributions.
- Troubleshoot OS boot, kernel, driver, filesystem, package and service-related issues.
- Solid knowledge of:
- LVM, RAID, XFS/EXT4
- NFS/SMB
- SSH
- DNS/DHCP
- NTP
- Systemd
- GRUB/Dracut
- SELinux
- Firewall management
- Analyze system logs, kernel messages, performance statistics and crash information.
- Troubleshoot OS-level performance, memory, CPU, disk I/O and network issues.
3. Application & Workload Support
- Troubleshoot infrastructure/application issues at the OS, middleware and application level.
- Monitor application and system performance and identify resource bottlenecks.
- Support installation, configuration and upgrades of infrastructure applications.
- Collect logs, statistics and diagnostic information for problem resolution.
- Coordinate with application/software vendors for unresolved issues.
4. AI / HPC / GPU Infrastructure
- Basic to intermediate knowledge of AI/ML and HPC infrastructure.
- Support GPU-based servers and AI workloads.
- Troubleshoot GPU detection, GPU driver, CUDA, PCIe and GPU performance-related issues.
- Knowledge of NVIDIA GPU architecture, NVIDIA drivers and CUDA will be an advantage.
- Exposure to HPC environments such as Slurm, OpenPBS, Kubernetes or similar platforms is desirable.
- Understanding of compute, storage and high-speed networking requirements for AI/HPC workloads.
5. Networking
- Basic to intermediate knowledge of enterprise networking.
- Troubleshoot Ethernet, VLAN, bonding/teaming, LACP, NIC and IP connectivity issues.
- Working knowledge of switches, routers and firewalls.
- Basic understanding of TCP/IP, DNS, DHCP, routing, VLAN and subnetting.
- Ability to perform network troubleshooting using tools such as:
- ping
- traceroute
- ip/ifconfig
- ethtool
- tcpdump
- iperf/iperf3
- Coordinate with network teams for switch, firewall and connectivity issues.
6. Cloud & Virtualization
- Basic understanding of cloud infrastructure and virtualization concepts.
- Exposure to VMware, KVM or other virtualization platforms.
- Understanding of VM compute, storage and networking requirements.
- Knowledge of cloud platforms such as AWS/Azure/GCP will be an added advantage.
7. Monitoring & Performance Management
- Monitor server health, CPU, memory, disk, network and application performance.
- Collect and analyze system statistics for capacity and performance analysis.
- Identify performance degradation and proactively highlight potential failures.
- Maintain infrastructure health and operational reports.
8. Incident & Problem Management
- Troubleshoot incidents from hardware firmware OS network application layers.
- Perform structured RCA (Root Cause Analysis) for recurring or critical incidents.
- Maintain incident, problem and change records.
- Follow defined SLA, escalation and change-management procedures.
- Coordinate with OEMs, vendors and internal technical teams for complex issues.
- Ensure proper documentation of troubleshooting activities and resolutions.
Technical Skills Mandatory
- Linux Server Administration
- Server hardware troubleshooting
- Hardware replacement and installation
- RAID / Storage fundamentals
- BIOS / BMC / Firmware
- IPMI / iLO / iDRAC
- OS installation and troubleshooting
- System performance monitoring
- TCP/IP networking fundamentals
- Switch and firewall fundamentals
- Shell scripting fundamentals
- Log analysis and troubleshooting
- Incident and problem management
Good to Have
- NVIDIA GPU / CUDA
- AI/ML infrastructure
- HPC environments
- Slurm / OpenPBS
- Kubernetes
- VMware / KVM
- SAN / NAS / NFS
- High-speed networking 10/25/40/100/200/400GbE
- InfiniBand
- Cloud AWS / Azure / GCP
- Python or advanced Shell scripting
- Automation using Ansible
Overall Competency The candidate should have the ability to troubleshoot an enterprise server end-to-end, starting from: Hardware Firmware BIOS/BMC RAID/Storage Network Operating System Drivers Application AI/HPC Workload
📌 Senior Engineer- HPC (Faridabad)
🏢 Netweb Technologies India
📍 Faridabad