15 Sep
|
Esconet Technologies
|
Delhi
15 Sep
Esconet Technologies
Delhi
Company : Esconet Technologies Location : Delhi/NCR – Okhla Phase 1, New Delhi (On-site) Experience : 5 Years About the Role Esconet Technologies is looking for a highly motivated HPC System Integrator/Administrator with a solid passion for cluster administration, system integration, and validation of HPC clusters. In this role, you will influence the overall integration, delivery, and management of largely open-source HPC products and solutions — spanning Intel and AMD processors, NVIDIA GPUs, storage, InfiniBand, and Linux software. A solid understanding of parallel processing (problem decomposition and work distribution), parallel programming (MPI, OpenMP), and computer architecture is essential for this role. Minimum Qualifications • Bachelor's degree in Computer Science, Computer Engineering, Computational Science, equivalent mathematical sciences, or a related field, with 3 years of relevant work experience; or an equivalent combination of education, training, and experience • 3 years of experience with software development in Linux • 3 years of experience with HPC clusters and systems integration • Ability to manage the AI stack up to the framework level • Experience with GPU cluster capabilities and storage integration (LUSTRE/BeeGFS, cluster storage) • Working knowledge of object-based storage, networking, InfiniBand, solution designing, and technical documentation Key Responsibilities • Install, configure, fine-tune, and troubleshoot multi-vendor, multi-site Linux HPC servers • Build and deploy open-source software as well as vendor/partner software • Diagnose and resolve system operational issues quickly and effectively • Verify full operation of systems, including network, systems,
and storage performance • Configure scheduling and queuing systems • Assist technical support teams with questions and issues encountered by customers • Coordinate with vendors to resolve hardware and software problems • Document system administration procedures for routine and complex tasks (wikis) • Maintain and monitor the security of HPC systems and servers • Design and configure HPC/Kubernetes clusters with NVIDIA GPU support, including MIG (Multi-Instance GPU) and AI Factory deployments • Deploy and manage cluster file systems such as Lustre, including OSS (Object Storage Server) and MDS (Metadata Server) node configuration • Build and maintain effective working relationships with coworkers, managers, and clients • Travel as needed for on-site cluster installation or maintenance (limited) Desired Skills & Experience • Building, configuring, and administering Linux distributions — Rocky Linux, Ubuntu, CentOS, RHEL, and SUSE • Windows Server administration • Expert knowledge of parallel/distributed file systems such as Lustre or IBM GPFS • Strong knowledge of networking and cluster-based distributed computing, including InfiniBand switch configuration • Experience deploying open-source and commercial HPC platforms • HPC cluster architecture design and configuration; AI cluster and Kubernetes cluster experience • Strong scripting skills — Bash, Python, Perl; Ansible for automation is a plus • Experience building dashboards (e.g., Flask-based) for cluster monitoring is a plus • Skilled in diagnosing and debugging complex HPC hardware/software issues, with proposed workarounds • Strong troubleshooting and root-cause analysis capabilities Key Skill Areas at a Glance Cluster Application & Design: HPC Cluster Application, HPC Cluster Architecture Designing, HPC Cluster Configuration AI/GPU Infrastructure: AI Cluster, Kubernetes Cluster, NVIDIA GPU, MIG, AI Factory Storage: Cluster File System (Lustre), Cluster Storage (OSS & MDS Nodes), BeeGFS, Object-Based Storage Networking: InfiniBand, InfiniBand Switch, Networking & Solution Design Automation/Scripting: Python, Bash, Perl, Ansible, Flask Dashboards OS Administration: Rocky Linux, Ubuntu, CentOS, RHEL, SUSE, Windows Server
📌 HPC Engineer (Delhi)
🏢 Esconet Technologies
📍 Delhi